yama6a / php-glyph-ocr
Pure PHP OCR for clean rendered text bitmaps such as image subtitles. A port of the nOCR engine from Subtitle Edit.
Requires
- php: ^8.2
- ext-zlib: *
Requires (Dev)
- ext-mbstring: *
- phpunit/phpunit: ^11.5
Suggests
- ext-gd: Needed only for Image::fromGd()
Provides
None
Conflicts
None
Replaces
None
This package is auto-updated.
Last update: 2026-10-03 01:54:48 UTC
README
A pure PHP OCR engine for clean rendered text bitmaps, such as the images of Blu-ray and DVD subtitles. It is a port of the nOCR engine from Subtitle Edit by Nikolaj Olsson.
The engine cuts an image into lines and glyphs and compares each glyph with a database of known glyphs. It reads text that a computer drew, in a font that the database knows. It does not read photos, scans or handwriting.
Install
Needs PHP 8.2 or later with ext-zlib. ext-gd is optional and only needed for Image::fromGd().
composer require yama6a/php-glyph-ocr
Usage
use GlyphOcr\GlyphDatabase; use GlyphOcr\Image; use GlyphOcr\Recognizer; $recognizer = new Recognizer(GlyphDatabase::subtitleFonts()); $result = $recognizer->recognize(Image::fromPng(file_get_contents('subtitle.png'))); $result->text(); // all lines, joined with "\n" $result->confidence(); // mean confidence of all glyphs, from 0 to 1 $result->lines[0]->text; // the first line $result->lines[0]->confidence; // mean confidence of the glyphs on the first line $result->lines[0]->chars[0]; // RecognizedChar: text, confidence, italic, glyph, sample
- Images:
Image::fromPng($bytes)decodes every PNG colour type without GD.Image::fromGd($gdImage)copies a GD image.Image::fromRgba($pixels, $width, $height)takes raw RGBA bytes or a list of integers. - Confidence: the share of the database glyph's line points that agree with the image, from 0 to 1. An unknown glyph reads as
*with confidence 0. - Ink: a pixel is ink when the sum of its premultiplied red, green and blue is at least
inkThreshold(200). So white or yellow text counts and a black outline does not. - State: a recognizer learns the glyph heights from the images it reads and uses them to split the next images. Use one recognizer per subtitle stream, or call
reset().
| Option | Default | Meaning |
|---|---|---|
inkThreshold |
200 |
Minimum premultiplied red + green + blue of an ink pixel, 1 to 765 |
spaceWidth |
null |
Empty columns that make a space. null uses a third of the typical glyph height of the line |
maxWrongPixels |
25 |
Error budget of the loose match passes |
fixLatinCase |
true |
Picks upper or lower case for letters such as o and O from their height |
unknownText |
'*' |
Text of a glyph that matches nothing |
italicSlant |
0.0 |
Above 0, a glyph that matches nothing is slanted back by this factor and tried again |
rightToLeft |
false |
Puts the glyphs of each line in right to left order |
minLineHeight |
12 |
Minimum line height in pixels until the recognizer has learned the glyph heights |
lineContext |
true |
Compares each glyph with the other glyphs of its line to pick I or l, o, O or 0, and comma or apostrophe. false keeps the database text, as Subtitle Edit does |
Databases and training
GlyphDatabase reads and writes the .nocr files of Subtitle Edit, version 1 and 2. The package ships two databases:
| Method | Glyphs | File size | Load time, PHP 8.5 | Memory |
|---|---|---|---|---|
GlyphDatabase::subtitleFonts() |
2,007 | 254 KB + 477 KB | 90 ms | 76 MB |
GlyphDatabase::latin() |
699 | 477 KB | 58 ms | 52 MB |
subtitleFonts(): glyphs trained on DejaVu Sans, Liberation Sans and Noto Sans, upright and italic, at 28 and 48 px, followed by the Latin database for other fonts. Liberation Sans has the metrics of Arial. Use this database for subtitles.latin(): the Latin database of Subtitle Edit, unchanged. Use it to get the same result as Subtitle Edit.- Other scripts and other fonts need their own database.
tests/fixtures/generator/train.php writes resources/SubtitleFonts.nocr again. The fonts and their licenses are in tests/fixtures/generator/fonts. The database holds line segments measured on rendered glyphs, and no font outlines.
Train a glyph from a sample that a person confirmed. The new glyph goes first, so it wins over older glyphs that match equally well:
use GlyphOcr\Trainer; $database = GlyphDatabase::fromFile('my-font.nocr'); $trainer = new Trainer(); foreach ($recognizer->recognize($image)->unknownChars() as $char) { $database->add($trainer->train($char->sample, askAPerson($char->sample->toAscii()))); } $database->save('my-font.nocr');
Recognizer::split($image)returns the glyph samples of each line without matching them, withnullfor a space.GlyphSample::merge([$left, $right])joins the parts of a character that the splitter cuts apart, such as"or%.- The trainer draws random line segments with a seeded generator, so the same sample and seed give the same glyph.
Accuracy
The golden images in tests/fixtures are white or yellow subtitles, 20 to 60 pixels, 1 or 2 lines. tests/fixtures/SOURCES.md describes them. One recognizer reads each set in order.
Character accuracy, with the lines read without error in brackets:
| Set | Fonts | subtitleFonts() |
latin() |
After training |
|---|---|---|---|---|
| Blu-ray, smooth RGBA | DejaVu Sans, Liberation Sans | 98.8% (9 of 11) | 93.8% (3 of 11) | 99.2% (10 of 11) |
| PGS palette | DejaVu Sans, Liberation Sans | 100% (8 of 8) | 93.0% (3 of 8) | 100% (8 of 8) |
| DVD, 4 colours, 24 to 30 px | DejaVu Sans, Liberation Sans | 95.5% (4 of 8) | 63.2% (0 of 8) | 94.2% (4 of 8) |
| Italic | DejaVu Sans, Liberation Sans | 98.0% (3 of 5) | 92.1% (1 of 5) | 97.0% (2 of 5) |
| No outline | DejaVu Sans, Liberation Sans | 98.9% (4 of 5) | 79.8% (1 of 5) | 96.6% (3 of 5) |
| Small, 20 to 24 px | DejaVu Sans, Liberation Sans | 98.1% (3 of 4) | 50.0% (0 of 4) | 92.3% (1 of 4) |
| Capital I and lower case l | DejaVu Sans, Liberation Sans, Noto Sans, Open Sans | 99.4% (15 of 17) | 90.2% (3 of 17) | 97.9% (10 of 17) |
| Unseen font | Open Sans | 96.7% (5 of 8) | 84.9% (0 of 8) | 96.7% (4 of 8) |
| All | 98.3% (51 of 66) | 85.0% (11 of 66) | 97.4% (42 of 66) |
- Character accuracy is 1 minus the edit distance divided by the length of the expected text.
- The Latin database has no glyphs of these fonts. The numbers show how it does on fonts it has not seen. The subtitle fonts database has no glyphs of Open Sans.
- "After training" adds to the Latin database one trained glyph per character for each font, size and style: the alphabet, digits and
.,!?'-:, drawn on a separate image. A new recognizer reads each image. - With
lineContext: false, the Latin database reads 80.3% of the characters. Then 37 of the 269 capital I and lower case l in the golden texts come out as the other letter. The subtitle fonts database reads 91.2%, and 63 come out as the other letter. WithlineContext: true, none do. - With
italicSlant: 0.2, the italic set reads 96.0% of the characters. - With the same database and
lineContext: false, the port reads every golden image exactly as Subtitle Edit does.SubtitleEditParityTestchecks this.
Speed
Mean time per golden image, one image of 1 or 2 lines, after the database is loaded:
| PHP 8.2 | PHP 8.5 | |
|---|---|---|
subtitleFonts() |
122 ms | 115 ms |
latin() |
244 ms | 247 ms |
Load subtitleFonts() |
184 ms | 90 ms |
Load latin() |
122 ms | 58 ms |
A full 1920x1080 frame with the same 2 lines takes about 1 second and 103 MB, so crop the image to the text where you can. A glyph that matches nothing is the slow case, because it runs through every match pass. So the subtitle fonts database is faster on the fonts it knows. Measured on one core of an x86_64 machine, without JIT.
Limits
- One text colour on a transparent or dark background. The ink threshold removes the outline, so text with a light outline or a light background does not work.
- The matcher compares shapes. Most sans-serif fonts draw I and l as the same bar, and l is about 5% taller. The recognizer picks the letter from the tops of the capitals, the ascenders and the other bars on the line. A line without them falls back to the heights of earlier lines, then to the neighbouring letters.
- No dictionary and no language model. Subtitle Edit fixes common OCR errors in a separate step, which this package does not port.
- Italic text needs italic glyphs in the database, or
italicSlant.
Exceptions
Every exception implements GlyphOcr\Exceptions\GlyphOcrException: InvalidArgumentException, InvalidImageException for a PNG that does not decode, and InvalidDatabaseException for a .nocr file that does not load.
Attribution
This package is a port of the nOCR engine from Subtitle Edit by Nikolaj Olsson, MIT license, at commit b8da12a. LICENSE keeps both copyright notices.
| This package | Subtitle Edit file in src/libuilogic/Ocr |
|---|---|
Internal/Matcher.php, Glyph.php |
NOcrDb.cs, NOcrChar.cs |
GlyphDatabase.php |
NOcrDb.cs, NOcrChar.cs (file format) |
Internal/Splitter.php, Internal/SplitItem.php |
NikseBitmapImageSplitter2.cs, ImageSplitterItem2.cs |
Internal/Bitmap.php |
NikseBitmap2.cs |
Internal/LinePoints.php |
NOcrLine.cs |
Internal/CaseFixer.php |
NOcrCaseFixer.cs |
Internal/LineHeightTracker.php |
OcrLineHeightTracker.cs |
Trainer.php |
NOcrChar.cs, NOcrLineGenerator.cs |
GlyphSample.php (merge) |
ExpandedOcrGroup.cs |
resources/Latin.nocr |
Ocr/Latin.nocr at the repository root, unchanged |
resources/SubtitleFonts.nocr is trained on renderings of these fonts:
| Font | License |
|---|---|
| DejaVu Sans 2.37 | Bitstream Vera license with public domain changes |
| Liberation Sans 2.1.5 | SIL Open Font License 1.1 |
| Noto Sans 2.015 | SIL Open Font License 1.1 |