belisoful / prado-bayesian
Bayesian classification and recommendation for the PRADO PHP framework: Naive Bayes spam filtering, multinomial/Bernoulli/complement Naive Bayes, TF-IDF weighting, evaluation metrics, and a probabilistic recommender.
Package info
github.com/belisoful/prado-bayesian
Type:prado4-extension
pkg:composer/belisoful/prado-bayesian
Requires
- php: >=8.1.0
- ext-mbstring: *
- pradosoft/prado: ^4.4@dev
Requires (Dev)
- friendsofphp/php-cs-fixer: ^3.94
- phpdocumentor/shim: ^3
- phpstan/phpstan: ^2.1
- phpunit/phpunit: ^10
Suggests
Provides
None
Conflicts
None
Replaces
None
README
Bayesian classification and recommendation for the PRADO PHP Framework (version 4.4+), implemented as a PRADO 4 extension.
Pre-release (0.x). This extension targets the PRADO
masterbranch (the upcoming 4.4 release), which adds theextra.prado.bootstrap/error-messages/class-mapComposer plugin hooks it relies on. It does not work with PRADO 4.3.x; because it depends on PRADO'smasterbranch (^4.4@dev), your application needsminimum-stability: dev(see Installation). See Versioning and compatibility for what may change before 1.0.0.
The module is designed for three common use cases out of the box:
- Spam filtering — train a Naive Bayes classifier on labeled text and classify new documents, getting the winning category and a score per category.
- Tagging — train documents under any number of labels and get an independent probability for every label (
phpandsecurity, or nothing). - Recommendation — score items for a user from observed item/category interactions, ranking by the score of the "likes this" class.
Training is incremental and reversible: trainOne() adds a document, untrainOne() withdraws it exactly. The raw scores are normalized Naive Bayes log-posteriors — a ranking that sums to one, not calibrated probabilities, because Naive Bayes is overconfident by construction. Fit a calibration on held-out documents (calibrate(), temperature scaling for a classifier and Platt scaling per label for a tagger) and the scores become probability estimates that are saved with the model; see Concepts.
The classifier, tokenizer, and storage are decoupled, so swapping in a different model family, token strategy, or persistence layer is a one-line configuration change.
Documentation
This README is the quick start. Deeper material lives in docs/:
| Page | What it covers |
|---|---|
| Concepts | The pipeline, the three Naive Bayes event models, smoothing, TF-IDF, log-space arithmetic, tokenization, and evaluation |
| Class reference | Every public class and interface by namespace, with its role and public API |
| Storage backends | The IBayesianStorage contract, the four backends, and how to choose |
| Configuration | Module and service wiring, and the full error-code list |
Requirements
| Requirement | Scope | Purpose |
|---|---|---|
| PHP 8.1 to 8.5 | required | Language runtime (each version is in CI; 8.4 and 8.5 run deprecation-free) |
ext-mbstring |
required | Multibyte-safe tokenization (every tokenizer uses mb_*) |
PRADO Framework ^4.4@dev |
required (Composer installs it) | TComponent, TService, TModule, TDbPropertiesTrait, and the extra.prado.* Composer plugin hooks |
ext-pdo |
suggested | Required by TSqlBayesianStorage (via Prado's TDbConnection) for SQL-backed persistence |
ext-redis |
suggested | Required by TRedisBayesianStorage for Redis-backed persistence |
SQL and Redis are opt-in. Add the extension you need with:
composer require ext-pdo # for TSqlBayesianStorage composer require ext-redis # for TRedisBayesianStorage
TMemoryBayesianStorage (default) and TFileBayesianStorage need no extension. Configuring TSqlBayesianStorage or TRedisBayesianStorage without the matching PHP extension is a configuration error (bayesian_storage_pdo_missing / bayesian_storage_redis_missing); there is no silent fallback.
Installation
composer require belisoful/prado-bayesian
The framework is a real requirement of this package — pradosoft/prado at ^4.4@dev — so
Composer installs it for you. Packagist already serves PRADO's master branch as dev-master,
which the framework aliases to 4.4.x-dev, so ^4.4@dev resolves from Packagist with no extra
repository. Your application needs only the two stability flags and the asset-packagist
repository:
{
"repositories": [
{ "type": "composer", "url": "https://asset-packagist.org" }
],
"require": {
"belisoful/prado-bayesian": "^0.1"
},
"minimum-stability": "dev",
"prefer-stable": true
}
minimum-stability: dev with prefer-stable: true is what lets ^4.4@dev resolve while every
other dependency still comes from its stable release. The constraint matches PRADO's
dev-master through the branch alias (dev-master → 4.4.x-dev) declared in the framework's
own composer.json.
The asset-packagist repository is the framework's requirement, not this package's — PRADO
depends on bower-asset/* packages, and because Composer reads repositories only from the
root project, never from a dependency, your application has to list it; installing without it
fails with bower-asset/jquery ... could not be found.
Name pradosoft/prado in your own require as well if you want to pin the framework version
your application runs; nothing here prevents it.
The package's config/ folder holds what PRADO's third-party plugin support reads system-wide from composer.json extra.prado: errorMessages.txt (the bayesian_* exception codes, registered through error-messages) and prado-bayesian-classes.json (Prado3-style short class names → PHP FQNs, registered through class-map, so TNaiveBayesClassifier resolves in Prado3-style configuration). Both load for every installed extension whether or not the bootstrap module is used.
What it provides
| Class | Namespace | Role |
|---|---|---|
TBayesianModule |
Belisoful\Prado\Util\Bayesian |
The extra.prado.bootstrap module; owns the configured default classifier |
TBayesianService |
Belisoful\Prado\Web\Services |
A TService exposing classification and recommendation over the PRADO service pipeline (HTTP request); opt-in access control through PRADO authorization rules and permissions |
IBayesianClassifier |
Belisoful\Prado\Util\Bayesian\Classifier |
The classifier contract: train(), trainOne(), untrain(), untrainOne(), classify(), score(), save(), load() |
TNaiveBayesClassifier |
Belisoful\Prado\Util\Bayesian\Classifier |
The classic Naive Bayes (multinomial event model with Laplace smoothing) — the default spam filter |
TMultinomialNaiveBayes |
Belisoful\Prado\Util\Bayesian\Classifier |
Multinomial Naive Bayes; counts token occurrences per category |
TBernoulliNaiveBayes |
Belisoful\Prado\Util\Bayesian\Classifier |
Bernoulli Naive Bayes; tracks token presence/absence per document |
TComplementNaiveBayes |
Belisoful\Prado\Util\Bayesian\Classifier |
Complement Naive Bayes; well-suited to imbalanced text classification |
IBayesianTokenizer |
Belisoful\Prado\Util\Bayesian\Tokenizer |
The tokenizer seam: input text → list of feature tokens |
TWordTokenizer |
Belisoful\Prado\Util\Bayesian\Tokenizer |
Default word tokenizer; lowercases, strips punctuation, drops short tokens, supports stop words |
TNGramTokenizer |
Belisoful\Prado\Util\Bayesian\Tokenizer |
Character or word n-grams (n configurable) |
TRegexTokenizer |
Belisoful\Prado\Util\Bayesian\Tokenizer |
A regex-driven tokenizer for custom patterns |
TBayesianTokenizerChain |
Belisoful\Prado\Util\Bayesian\Tokenizer |
Composes multiple tokenizers (each contributes its tokens) |
TBayesianTokenizerTrait |
Belisoful\Prado\Util\Bayesian\Tokenizer |
Shared tokenizer plumbing: property-driven exportConfig()/importConfig(), safe matchAll(), normalizeText() |
TBayesianTokenizerFactory |
Belisoful\Prado\Util\Bayesian\Tokenizer |
Serializes/restores tokenizers into the saved model, UTF-8 scrubbing, regex validation |
IBayesianVocabulary |
Belisoful\Prado\Util\Bayesian |
The vocabulary seam: resident, or read per token from storage. getVocabulary() returns this |
TBayesianVocabulary |
Belisoful\Prado\Util\Bayesian |
The resident vocabulary; per-category token counts, totals, smoothing |
TLazyBayesianVocabulary |
Belisoful\Prado\Util\Bayesian |
The storage-backed vocabulary; reads a document's tokens per classification via IBayesianTokenStorage |
TLazyBayesianCategory |
Belisoful\Prado\Util\Bayesian |
A category whose per-token counts come from the vocabulary's last prefetch |
TBayesianCategory |
Belisoful\Prado\Util\Bayesian |
One category: its name, document count, and token counts |
TBayesianTrainingSet |
Belisoful\Prado\Util\Bayesian |
An iterable labeled training set: maps categories to tokenized documents |
TBayesianModelConverter |
Belisoful\Prado\Util\Bayesian |
Rewrites a whole-payload model into a per-token backend without retraining |
IBayesianTagger / TBayesianTagger |
Belisoful\Prado\Util\Bayesian |
Multi-label tagging: one-versus-rest Naive Bayes over one shared model, an independent probability per label |
TBayesianPayload |
Belisoful\Prado\Util\Bayesian |
Typed reads out of decoded payloads and configuration arrays |
TBayesianTokenHistogram |
Belisoful\Prado\Util\Bayesian |
The "count of counts" histograms Bernoulli and Complement sum over instead of walking the vocabulary |
TTemperatureScaling |
Belisoful\Prado\Util\Bayesian\Calibration |
Calibrates a classifier's scores into probabilities with one fitted temperature |
TPlattScaling |
Belisoful\Prado\Util\Bayesian\Calibration |
Calibrates a binary decision value (a tagger's per-label log-odds) with a fitted logistic curve |
TFIdf |
Belisoful\Prado\Util\Bayesian\Math |
Term-frequency × inverse-document-frequency weighting |
TBayesMath |
Belisoful\Prado\Util\Bayesian\Math |
Log-space arithmetic helpers used by the classifiers to avoid underflow |
TConfusionMatrix |
Belisoful\Prado\Util\Bayesian\Evaluation |
Confusion matrix for evaluating a classifier against a labeled set |
TBayesianMetrics |
Belisoful\Prado\Util\Bayesian\Evaluation |
Precision, recall, F1, accuracy, macro/micro averages |
TCalibrationMetrics |
Belisoful\Prado\Util\Bayesian\Evaluation |
Log loss, Brier score and expected calibration error, to judge a calibration on held-out data |
IBayesianStorage |
Belisoful\Prado\Util\Bayesian\Storage |
The persistence seam for a trained model |
IBayesianTokenStorage |
Belisoful\Prado\Util\Bayesian\Storage |
A storage backend that also serves a model per token, for models larger than a process |
IBayesianHistogramStorage |
Belisoful\Prado\Util\Bayesian\Storage |
A per-token backend that also keeps the token histograms, so Bernoulli and Complement train incrementally against it |
TMemoryBayesianStorage |
Belisoful\Prado\Util\Bayesian\Storage |
Process-local in-memory storage (default; no I/O) |
TFileBayesianStorage |
Belisoful\Prado\Util\Bayesian\Storage |
JSON file storage (good for development, small models, single host, single writer) |
TSqlBayesianStorage |
Belisoful\Prado\Util\Bayesian\Storage |
SQL-backed storage via TDbConnection (SQLite, MySQL, PostgreSQL); whole-payload or per-token (Mode); per-token training is an atomic increment, safe for concurrent writers; connection through TDbPropertiesTrait |
TRedisBayesianStorage |
Belisoful\Prado\Util\Bayesian\Storage |
Redis-backed storage for shared hosts; whole-payload or per-token (Mode); per-token training is a Lua script of HINCRBY increments, safe for concurrent writers (requires ext-redis) |
IBayesianRecommender |
Belisoful\Prado\Util\Bayesian |
The recommender contract: recommend() for a user/item context |
TBayesianRecommender |
Belisoful\Prado\Util\Bayesian |
A probabilistic recommender built on top of any IBayesianClassifier |
Architecture
entry points TBayesianModule ────────────► TBayesianService
(extra.prado.bootstrap; (TService; HTTP
owns default classifier) classify/recommend)
│ both resolve
▼
seam IBayesianClassifier ◄──────► IBayesianStorage
│ memory / file / SQL / Redis
│ implemented by
▼
classifiers TNaiveBayesClassifier TBayesianRecommender
▲ extends (ranks candidates by
┌───────────────┼───────────────┐ P(positive), reusing
TMultinomialNaiveBayes TBernoulli- TComplement- any classifier)
NaiveBayes NaiveBayes
│ reads / writes
▼
training state TBayesianVocabulary ─── TBayesianCategory ─── TBayesianTrainingSet
│ scores with
▼
math TBayesMath ─── TFIdf
│ features from
▼
tokenizers IBayesianTokenizer
TWordTokenizer / TNGramTokenizer / TRegexTokenizer / TBayesianTokenizerChain
│
(text in, tokens out)
The layers stack cleanly:
- Math —
TBayesMathworks in log-space, so the Naive Bayes product of thousands of small probabilities never underflows.TFIdfweights token contributions by how discriminating they are across the corpus. - Tokenizer —
IBayesianTokenizeris the seam between text and features. DefaultTWordTokenizeris good enough for spam filtering; swap inTNGramTokenizerfor language-agnostic content orTRegexTokenizerfor structured input. - Vocabulary & categories —
IBayesianVocabularyis the statistics the classifier scores against, behind an interface so they need not all be resident:TBayesianVocabularyholds the whole model,TLazyBayesianVocabularyreads a document's tokens from storage per classification.TBayesianCategoryrepresents one class.TBayesianTrainingSetis the labeled corpus in training-time form. - Classifiers — All implement
IBayesianClassifierand accept any tokenizer + storage.TNaiveBayesClassifieris the canonical spam filter and the base class of the other three;TMultinomialNaiveBayes,TBernoulliNaiveBayes, andTComplementNaiveBayesoverride only the likelihood, so switching event model is a one-line change. Each writes a distinctkindmarker into its saved model, so several variants can share one storage backend safely. - Storage —
IBayesianStoragepersists a trained model.TMemoryBayesianStorageis the no-I/O default;TFileBayesianStoragewrites JSON;TSqlBayesianStorageuses Prado'sTDbConnection/TDbCommandfor SQL-backed persistence (SQLite, MySQL, PostgreSQL), configured throughTDbPropertiesTraitlike any other Prado database component, and can store a model per token (Mode="token") so it is bounded by the database rather than by PHP memory;TRedisBayesianStoragescales across processes and hosts via Redis, and like the SQL backend can store a model per token (Mode="token"), though there the model lives in Redis's RAM rather than on disk. Whole-payload storage is single-writer (a save replaces the model); per-token storage is multi-writer (training is an atomic increment), which is what a model trained from concurrent web requests and workers needs. See Storage → Concurrency. - Recommender —
TBayesianRecommenderreuses the classifier: train it on user/item interactions with a positive and a negative label (PositiveCategorydefaults toliked), then ask it to rank candidate items. - Module & service —
TBayesianModuleis theextra.prado.bootstrapentry point that owns the configured classifiers and storage; one module can hold several models over one backend.TBayesianServiceexposes a classifier and the recommender over the PRADO service pipeline (HTTP), sourcing its classifier from the module.
Usage
Spam filter (the default)
use Belisoful\Prado\Util\Bayesian\Classifier\TNaiveBayesClassifier; $classifier = new TNaiveBayesClassifier(); $classifier->setName('comment-spam'); foreach ([ 'Buy cheap watches now!!!', 'Limited time offer, click here', 'Congratulations, you have won a prize', ] as $document) { $classifier->trainOne('spam', $document); } foreach ([ 'Hey, are we still meeting for lunch tomorrow?', 'I attached the report you asked for.', 'Thanks for the help with the bug fix.', ] as $document) { $classifier->trainOne('ham', $document); } $label = $classifier->classify('FREE VIAGRA!!! Lowest prices online'); // 'spam' $spam = $classifier->isSpam('FREE VIAGRA!!! Lowest prices online'); // true $score = $classifier->score('Buy cheap watches now!!!'); // ['spam' => 0.99..., 'ham' => 0.00...]
Persist a trained model
use Belisoful\Prado\Util\Bayesian\Storage\TFileBayesianStorage; $storage = new TFileBayesianStorage(); $storage->setDirectory('/var/lib/myapp/bayesian'); $classifier->setStorage($storage); $classifier->save(); // serialize the trained model $classifier->load('comment-spam'); // restore on a future request
The saved state carries the tokenizer class and its settings, so a model trained with a TNGramTokenizer (or a TBayesianTokenizerChain) tokenizes identically after load() into a fresh classifier. Each classifier variant writes a kind marker and refuses to load a payload saved by a different variant (bayesian_classifier_kind_mismatch); load a model with the class that saved it. TSqlBayesianStorage creates its table on first use with driver-aware DDL (VARCHAR(191)/LONGTEXT on MySQL); set AutoCreateTable="false" to manage the schema yourself.
As a PRADO module
PRADO reads either XML or PHP application configuration; both forms are shown throughout.
protected/application.xml
<modules> <module id="bayesian" class="Belisoful\Prado\Util\Bayesian\TBayesianModule" DefaultClassifier="comment-spam"> <!-- optional: pick the classifier class and set its properties (default: TNaiveBayesClassifier) --> <classifier class="Belisoful\Prado\Util\Bayesian\Classifier\TComplementNaiveBayes" Alpha="0.5" /> <storage class="Belisoful\Prado\Util\Bayesian\Storage\TFileBayesianStorage" Directory="/var/lib/myapp/bayesian" /> </module> </modules> <services> <service id="bayesian" class="TBayesianService" ModuleID="bayesian" MaxTextLength="65536" /> </services>
protected/application.php
<?php return [ 'modules' => [ 'bayesian' => [ 'class' => 'Belisoful\Prado\Util\Bayesian\TBayesianModule', // Module properties go under 'properties'; the <classifier>/<storage> child // elements become sibling keys of 'class'. 'properties' => ['DefaultClassifier' => 'comment-spam'], 'classifier' => ['class' => 'TComplementNaiveBayes', 'Alpha' => 0.5], 'storage' => ['class' => 'TFileBayesianStorage', 'Directory' => '/var/lib/myapp/bayesian'], ], ], 'services' => [ 'bayesian' => [ 'class' => 'TBayesianService', 'properties' => ['ModuleID' => 'bayesian', 'MaxTextLength' => 65536], ], ], ];
The two are equivalent. Note the shape difference: an XML attribute on <module> is a module
property and lives under 'properties', while <classifier> and <storage> are child
elements and become their own keys alongside 'class'. Within those two, the class and its
properties sit side by side, exactly as the XML attributes do.
Registering by package name works in PHP too — but then omit 'class', because the class comes
from extra.prado.bootstrap and supplying both is a configuration error:
'modules' => [ 'belisoful/prado-bayesian' => [ 'properties' => ['DefaultClassifier' => 'comment-spam'], 'storage' => ['class' => 'TFileBayesianStorage', 'Directory' => '/var/lib/myapp/bayesian'], ], ],
Both forms of module registration work: <module id="belisoful/prado-bayesian"> (the package name; the class comes from extra.prado.bootstrap) or <module id="bayesian" class="Belisoful\Prado\Util\Bayesian\TBayesianModule">. Register the service by its class-map short name TBayesianService: PRADO master resolves <service class="…"> through Prado::usingClass(), which currently does not fall back to the Composer autoloader for a not-yet-loaded fully-qualified name outside the framework directory.
DefaultClassifier names the model: the classifier takes that name, and if the storage already holds a saved model of that name it is loaded when the module initializes. A model that has not been trained/saved yet is simply empty until you train and save() it; a storage that cannot be reached is a configuration error at startup, not a silent empty model.
Application code reaches the configured default classifier through the module:
$module = Prado::getApplication()->getModule('bayesian'); $label = $module->getClassifier()->classify($text);
TBayesianService is read-only over HTTP (no training or deletion) and answers with JSON. It enforces no access control by default — restrict it with PRADO authorization rules (an <authorization> element of the service) or a TPermissionsManager before exposing it; see Configuration → Access control. With the service id bayesian:
| Request | Response |
|---|---|
?bayesian&text=Free+pills (or &action=classify) |
{"category":"spam","scores":{"spam":0.98,"ham":0.02}} |
?bayesian&text=Free+pills&category=spam |
adds "isSpam": true |
?bayesian&action=recommend&context[]=red+shoes&candidates[]=red+hat&candidates[]=blue+hat |
{"scores":{"red hat":0.7,"blue hat":0.4}} (always a JSON object, highest first) |
?bayesian&action=tag&text=escape+the+query+parameters |
{"tags":{"php":0.91,"security":0.78},"calibrated":false} (labels above TagThreshold, highest first, at most MaxTags) |
Errors are JSON with an HTTP status: 400 {"error":"bayesian_service_text_required",...} for a missing/malformed parameter or unknown action, 401/403 when access is refused, 413 when text, the joined context or a candidate exceeds MaxTextLength (default 65536 bytes, 0 = unlimited) or more than MaxCandidates (default 100) candidates are sent, and 503 bayesian_classifier_not_trained when no model has been trained yet. ModuleID selects which TBayesianModule supplies the classifier (default: the first one registered).
Multiple models
A model is identified by its name, and storage is keyed by that name — so one storage backend holds as many models as you like. There is no registry to configure; naming and saving is all it takes.
$storage = new TFileBayesianStorage(); $storage->setDirectory('/var/lib/myapp/bayesian'); $spam = new TNaiveBayesClassifier(); $spam->setStorage($storage); $spam->setName('comment-spam'); $spam->trainOne('spam', 'cheap pills buy now'); $spam->trainOne('ham', 'project meeting tomorrow'); $spam->save(); $language = new TBernoulliNaiveBayes(); // a different variant, same storage $language->setStorage($storage); $language->setName('language-id'); $language->trainOne('en', 'the quick brown fox'); $language->trainOne('fr', 'le renard brun rapide'); $language->save(); $storage->list(); // ['comment-spam', 'language-id'] — sorted ascending
Two things make this safe to do in one store:
load()replaces the classifier's state entirely. You can reuse a single instance across models —$c->load('comment-spam')then$c->load('language-id')— and no categories, counts, or tokenizer settings survive from the previous model.- Variants cannot be crossed. Each classifier class writes a
kindmarker, so loading a Bernoulli model into aTComplementNaiveBayesthrowsbayesian_classifier_kind_mismatchrather than silently scoring with the wrong math. (TNaiveBayesClassifierandTMultinomialNaiveBayesare interchangeable — they are the same model.)
Whether to reuse one instance or hold several depends on access pattern. Holding several keeps
every model resident in memory at once; reusing one re-reads and re-decodes the payload on each
load(). See Model size and storage limits for what that costs.
The module's DefaultClassifier covers only the single model the application boots with. For
the rest, construct classifiers yourself and hand them the module's storage:
$storage = $this->getApplication()->getModule('bayesian')->getStorage();
Model size and storage limits
In the default payload mode, a backend loads the whole model. The model is one JSON unit —
one file, one row, one Redis key — and load() decodes all of it into PHP arrays before the
first classification. That mode is the right choice whenever the model fits comfortably in a
request and has one writer.
TSqlBayesianStorage and TRedisBayesianStorage also offer Mode="token", where the model is
stored per token and a classification reads only the document's own tokens — so a loaded model
costs kilobytes of PHP memory regardless of its size, loading it takes well under a millisecond
where the payload form takes tens of milliseconds and tens of megabytes, and training one
document writes that document's rows instead of re-serializing the model. The figures, and the
script that reproduces them on your machine (composer benchmark), are in
Storage backends.
The rules of thumb for payload mode: the payload runs 30–40 bytes per token-per-category,
because each category stores its own occurrence and document counts for every token it has seen,
plus one corpus-wide document-frequency map; and the decoded PHP structure is roughly 3–4× the
JSON, since PHP's hash tables cost far more per entry than the text does. Budget the sum of
both: json_decode() holds the string and the growing array at the same time.
Size therefore scales with vocabulary times categories, not vocabulary alone — ten categories
over the same words is five times the model of two. Trimming the vocabulary is the effective
lever: raise MinLength, supply StopWords, or prefer word tokens over character n-grams, which
generate far more distinct features.
| Backend | Ceiling on one model | Practical limit |
|---|---|---|
TMemoryBayesianStorage |
PHP memory_limit |
Also gone at end of process |
TFileBayesianStorage |
Filesystem | One .json file per model, read whole via file_get_contents() |
TSqlBayesianStorage |
LONGTEXT 4 GB (MySQL), TEXT ~1 GB (PostgreSQL/SQLite) |
MySQL max_allowed_packet (64 MB default) usually binds first |
TRedisBayesianStorage |
512 MB per value (payload); the Redis instance's RAM (token) | Redis holds the whole model in RAM either way |
In every case PHP's memory_limit binds long before the backend's own ceiling: a 64 MB payload
needs roughly 200–250 MB of PHP memory to decode and hold. If you need models larger than a
process can hold, the fix is a smaller feature space — not a different backend.
Untraining
Every training call has an exact inverse. A document that was trained and then untrained leaves no trace: the counts return to what they were, a token no document contains any more leaves the vocabulary, and a category left without documents disappears.
$classifier->trainOne('spam', 'cheap watches'); $classifier->untrainOne('spam', 'cheap watches'); // as if it had never been trained $classifier->untrain($trainingSet); // withdraws a whole set
Untraining works against every storage layout, including per-token models trained from many processes at once (the deltas are atomic decrements, clamped at zero). Pass the document exactly as it was trained — the same text through the same tokenizer, or the same pre-tokenized list — or the counts of the tokens that differ will be off by one.
Calibrated probabilities
score() normalizes the Naive Bayes log-posteriors, which is a ranking, not a probability: a
document that is 70% likely to be spam routinely scores 0.99. Fit a calibration on labeled
documents the model was not trained on, and score() returns probability estimates from then
on; the calibration is saved with the model.
$classifier->calibrate($heldOutSet); // fits a TTemperatureScaling and installs it $classifier->getIsCalibrated(); // true $classifier->score('cheap pills'); // now calibrated; classify() is unchanged $classifier->logScores('cheap pills'); // the raw log-posteriors, always available $classifier->getCalibration()->getTemperature();
Judge the result with TCalibrationMetrics on a third set of documents: distributionLogLoss()
and distributionCalibrationError() before and after. Refit after substantial further training.
Multi-label tagging
A TBayesianTagger trains a document under each of its labels and returns an independent
probability for every label, so a document can carry several tags or none. It is one-versus-rest
Naive Bayes over one shared model, so it costs one write per label plus one, and every storage
backend and both layouts work unchanged.
use Belisoful\Prado\Util\Bayesian\TBayesianTagger; $tagger = new TBayesianTagger(); $tagger->getClassifier()->setStorage($storage); $tagger->getClassifier()->setName('post-tags'); $tagger->train(['php', 'security'], 'Validate every request parameter before it reaches the SQL query'); $tagger->train(['cooking'], 'Simmer the sauce for twenty minutes'); $tagger->train([], 'The meeting moved to Tuesday'); // a negative example for every label $tagger->probabilities('Escape the query parameters'); // ['php' => 0.91, 'security' => 0.78, 'cooking' => 0.04] $tagger->tag('Escape the query parameters'); // ['php' => 0.91, 'security' => 0.78] — above Threshold, highest first $tagger->untrain(['cooking'], 'Simmer the sauce for twenty minutes'); $tagger->calibrate($heldOutExamples); // a TPlattScaling per label, saved with the model $tagger->getClassifier()->save();
Threshold (default 0.5) and MaxTags (default unlimited) shape tag(). The service exposes
the same as action=tag. The underlying classifier's own classify() is not meaningful for a
tagged model — score through the tagger.
Recommendation
Train a TBayesianRecommender's underlying classifier on user behavior: the "positive" category is the items the user engaged with; the "negative" category is items they ignored. Then ask it to rank candidates for a new user context.
use Belisoful\Prado\Util\Bayesian\TBayesianRecommender; $rec = new TBayesianRecommender(); $rec->setPositiveCategory('liked'); $classifier = $rec->getClassifier(); foreach (['red shoes', 'blue sneakers', 'leather boots'] as $item) { $classifier->trainOne('liked', $item); } foreach (['red hat', 'blue scarf', 'leather belt'] as $item) { $classifier->trainOne('ignored', $item); } $top = $rec->recommend(['red shoes', 'leather belt'], ['red sneakers', 'blue sneakers', 'red hat', 'leather wallet']); // ['red sneakers' => 0.83, 'blue sneakers' => 0.83, 'leather wallet' => 0.5, 'red hat' => 0.17] // (values from this exact toy corpus; ties keep candidate order)
Note that unseen words are ignored at classification time (they carry no learned evidence), so a candidate made only of tokens the classifier has never seen scores at the class prior.
// A purely numeric identifier such as '123' comes back as the integer key 123 (PHP array // semantics); json_encode it with JSON_FORCE_OBJECT if the result must stay a map.
Evaluate a model
use Belisoful\Prado\Util\Bayesian\Evaluation\TConfusionMatrix; use Belisoful\Prado\Util\Bayesian\Evaluation\TBayesianMetrics; $matrix = new TConfusionMatrix(['spam', 'ham']); foreach ($labeledTestSet as $document => $expectedLabel) { $predicted = $classifier->classify($document); $matrix->record($expectedLabel, $predicted); } $metrics = new TBayesianMetrics($matrix); $accuracy = $metrics->getAccuracy(); $spamF1 = $metrics->getF1('spam'); echo "accuracy={$accuracy}, spam-F1={$spamF1}\n";
Full lifecycle example
Every part of the extension in one script: a composed tokenizer, a persistent backend, a Complement Naive Bayes classifier, a labeled training set, save/reload, classification, evaluation on held-out documents, and incremental retraining.
use Belisoful\Prado\Util\Bayesian\Classifier\TComplementNaiveBayes; use Belisoful\Prado\Util\Bayesian\Evaluation\TBayesianMetrics; use Belisoful\Prado\Util\Bayesian\Evaluation\TConfusionMatrix; use Belisoful\Prado\Util\Bayesian\Storage\TFileBayesianStorage; use Belisoful\Prado\Util\Bayesian\TBayesianTrainingSet; use Belisoful\Prado\Util\Bayesian\Tokenizer\TBayesianTokenizerChain; use Belisoful\Prado\Util\Bayesian\Tokenizer\TNGramTokenizer; use Belisoful\Prado\Util\Bayesian\Tokenizer\TWordTokenizer; // 1. Tokenizer — words for meaning, character trigrams to survive obfuscation ("ch34p"). // Chain members each see the original text; their outputs are concatenated. $words = new TWordTokenizer(); $words->setMinLength(3); $words->setStopWords(['the', 'and', 'for', 'you', 'your']); $trigrams = new TNGramTokenizer(); $trigrams->setN(3); $trigrams->setCharacters(true); $tokenizer = new TBayesianTokenizerChain(); $tokenizer->addTokenizer($words); $tokenizer->addTokenizer($trigrams); // 2. Storage — one JSON file per model, written atomically. $storage = new TFileBayesianStorage(); $storage->setDirectory('/var/lib/myapp/bayesian'); // 3. Classifier — Complement NB handles the class imbalance real spam corpora have. $classifier = new TComplementNaiveBayes(); $classifier->setName('comment-spam'); // the storage key $classifier->setTokenizer($tokenizer); $classifier->setStorage($storage); $classifier->setAlpha(0.5); // must be > 0 $classifier->setUseTfidf(true); // 4. Train from a labeled set. $training = new TBayesianTrainingSet(); foreach ([ 'Buy cheap watches now, limited time offer', 'Congratulations! You have won a free prize', 'Lowest prices online, click here to order', 'Cheap pills delivered discreetly, order now', 'Make money fast working from home', ] as $doc) { $training->add('spam', $doc); } foreach ([ 'Are we still meeting for lunch tomorrow?', 'I attached the quarterly report you asked for', 'Thanks for the help with that bug fix', 'The deployment finished, staging looks healthy', 'Can you review my pull request this week?', ] as $doc) { $training->add('ham', $doc); } $classifier->train($training); // getIsTrained() === true, categories ["spam","ham"], 10 documents // 5. Persist. The payload carries the tokenizer class and its settings too. $classifier->save(); $storage->list(); // ['comment-spam'] — the file is ~30 KB for this toy corpus // 6. Reload in a later request, into a fresh instance that was never configured. $loaded = new TComplementNaiveBayes(); $loaded->setStorage($storage); $loaded->load('comment-spam'); // The chain and both its members come back, and Alpha is 0.5 again — nothing to re-set. // 7. Classify. $probe = 'Cheap watches, lowest prices, order now'; $label = $loaded->classify($probe); // 'spam' $scores = $loaded->score($probe); // ['spam' => 0.512, 'ham' => 0.488] // 8. Evaluate on documents the model was NOT trained on. $heldOut = [ ['Free prize waiting, claim it today', 'spam'], ['Order cheap pills online now', 'spam'], ['Lunch tomorrow at the usual place?', 'ham'], ['Please review the attached report', 'ham'], ]; $matrix = new TConfusionMatrix(['spam', 'ham']); foreach ($heldOut as [$text, $expected]) { $matrix->record($expected, $loaded->classify($text)); } $metrics = new TBayesianMetrics($matrix); $metrics->getAccuracy(); // 1.00 $metrics->getF1('spam'); // 1.00 $metrics->getMacroF1(); // 1.00 — prefer macro when classes are imbalanced // 9. Training is incremental: correct a mistake and re-save, no full retrain. $loaded->trainOne('spam', 'Exclusive offer just for you, act now'); $loaded->save();
The commented values are this script's actual output. Two things they illustrate honestly: a ten-document corpus produces a narrow margin (0.512 vs 0.488) even when the label is right, because Complement NB normalizes each category's weight vector and there is very little evidence to separate them — real corpora run to thousands of documents and separate far more sharply. And perfect scores on four held-out documents mean nothing statistically; they show the evaluation wiring works, not that the model is good.
Versioning and compatibility
The package follows Semantic Versioning and is pre-1.0. Until 1.0.0:
- a minor release (0.x → 0.y) may change public APIs, stored formats and defaults; every such change is listed in CHANGELOG.md with a migration note;
- a patch release changes behavior only to fix a bug, and never changes a stored format;
- stored models carry a
formatVersion(payloads) and alayoutVersion(per-token metadata), so a release that changes a format can upgrade or refuse an older model instead of misreading it, and a release always reads the previous version's models.
From 1.0.0 the usual guarantee applies: only a major release may break compatibility.
Development
composer.json requires pradosoft/prado at ^4.4@dev, which resolves straight from Packagist
through the branch alias in PRADO's own composer.json (dev-master → 4.4.x-dev) —
composer install needs no repository entry and no further setup. To develop against a local
PRADO checkout instead, add a path repository to your working copy and leave it uncommitted:
composer config repositories.prado --json \
'{"type":"path","url":"../prado.master","options":{"versions":{"pradosoft/prado":"4.4.x-dev"}}}'
composer update pradosoft/prado
composer install composer fulltest # the full check: lint, code style, static analysis (PHPStan level 3), unit tests composer unittest # tests only composer fix # apply the code style composer coverage # tests with a coverage report composer integration # Composer-extension install check composer benchmark # model-size and load-time figures per storage mode
See CONTRIBUTING.md for the backend environment variables and the contribution rules, and SECURITY.md for reporting vulnerabilities.
composer integration builds a throwaway consumer project that requires this package through
Composer and asserts what extra.prado promises: the error-message file resolves bayesian_*
codes, the class map resolves the Prado3 short names, the bootstrap module boots under its
package id, and a real request reaches TBayesianService. It expects the framework at
../prado.master; pass another path as the first argument.
The consumer project is removed when the run passes. A failing run keeps it — and prints where —
so the half-built install can be inspected; KEEP_WORK_DIR=1 keeps it after a passing run too. A
work directory passed as the second argument belongs to the caller and is never removed.
The committed composer.json carries no machine-specific paths and no framework repository entry: pradosoft/prado at ^4.4@dev resolves straight from Packagist, which is what CI installs from too. A full check before committing is, in order: php -l, php-cs-fixer, phpstan, phpunit — composer fulltest.
Tests cover the math (log-space arithmetic, TF-IDF), tokenizers (word, n-gram, regex, chain, factory round-trips), the classifiers (Naive Bayes and the three variants over tokenized corpora, training and untraining, save/load including the tokenizer, the format version and the calibration), the calibrations (temperature and Platt scaling, the calibration metrics), the tagger (independent per-label probabilities, thresholds, untraining, calibration, and per-token equivalence), the recommender, the storage backends (including the ascending-order list() contract exercised with out-of-order saves, and the framework database-connection contract of TSqlBayesianStorage), the module (one classifier and several sharing one backend, each with its own tokenizer), the service (including its authorization rules, permissions and input limits), and the metrics. Per-token storage is tested for score-equivalence with the whole-payload layout across every variant, in both SQL and Redis, along with incremental training, the layout upgrade from 0.1.0, negative deltas, and the TBayesianModelConverter; concurrent training is tested with parallel processes against SQLite, MySQL, PostgreSQL and Redis; the Redis field encoding is additionally proven in isolation so it does not rest on a live server being present; and a package test checks that every source class is in the class map and every error code raised is defined and documented.
SQL tests skip cleanly when ext-pdo/pdo_sqlite is unavailable; Redis tests skip when ext-redis is absent or no server listens on 127.0.0.1:6379; the MySQL and PostgreSQL suites run only when BAYESIAN_MYSQL_DSN / BAYESIAN_PGSQL_DSN name a reachable server:
BAYESIAN_PGSQL_DSN="pgsql:host=127.0.0.1;port=5432;dbname=bayesian_test" BAYESIAN_PGSQL_USER=postgres vendor/bin/phpunit --testsuite unit
CI runs PHP 8.1 through 8.5 (plus a lowest-dependencies leg; 8.4 and 8.5 have also been run locally with deprecations displayed) against Redis, MySQL, and PostgreSQL service containers with BAYESIAN_REQUIRE_BACKENDS=1, which turns each of those skips into a failure — so a green build means every backend really ran, rather than quietly skipping. It also runs weekly, because the package tracks PRADO's master branch.
Line coverage is ~87% locally with only SQLite available, and higher with Redis, MySQL, and PostgreSQL running (composer coverage, needs Xdebug or PCOV; CI enforces a 93% floor). The gap between the two figures is almost entirely the Redis backend, which cannot run a line without ext-redis.
Locally the uncovered remainder is, in order of size:
TRedisBayesianStorage(the largest block) — its constructor throws withoutext-redis, so no instance can exist on a machine that lacks the extension: the whole class, both payload and per-token, is unreachable locally. The one branch that is reachable there — the missing-extension guard — is covered bytestConstructorThrowsWhenRedisExtensionMissing, and the per-token field encoding is covered byTRedisTokenEncodingTest, which reaches the pure helpers by reflection without constructing the class. Everything else runs in CI.- The MySQL-only branches of
TSqlBayesianStorage— theON DUPLICATE KEY UPDATEupserts in both the payload and per-token paths, reachable only against a real MySQL server. - IO- and DB-failure guards — a
tempnam/file_put_contentsfailure inTFileBayesianStorage, a connect failure inTSqlBayesianStorage. Reaching these needs the filesystem or server to fail between the check and the write. - Defensive guards the public API cannot reach at all — a smoothing denominator of zero (a positive
Alpharules it out), a zero TF-IDF weight (the smoothed IDF is always ≥ 1), an empty-string split behind a caller that already returned early, and a pad whose target width is never below the input. These are unreachable by construction, not merely untested.
See CHANGELOG.md for release notes.
License
BSD-3-Clause. See LICENSE.