Search by

yama6a / php-glyph-ocr

yama6a

Pure PHP OCR for clean rendered text bitmaps such as image subtitles. A port of the nOCR engine from Subtitle Edit.

Package info

github.com/yama6a/php-glyph-ocr

pkg:composer/yama6a/php-glyph-ocr

Statistics

Installs: 195

Dependents: 1

Suggesters: 1

Stars: 0

Open Issues: 0

0.1.1 2026-10-02 20:31 UTC

This package is auto-updated.

Last update: 2026-10-02 20:33:43 UTC


README

A pure PHP OCR engine for clean rendered text bitmaps, such as the images of Blu-ray and DVD subtitles. It is a port of the nOCR engine from Subtitle Edit by Nikolaj Olsson.

The engine cuts an image into lines and glyphs and compares each glyph with a database of known glyphs. It reads text that a computer drew, in a font that the database knows. It does not read photos, scans or handwriting.

Install

Needs PHP 8.2 or later with ext-zlib. ext-gd is optional and only needed for Image::fromGd().

composer require yama6a/php-glyph-ocr

Usage

use GlyphOcr\GlyphDatabase;
use GlyphOcr\Image;
use GlyphOcr\Recognizer;

$recognizer = new Recognizer(GlyphDatabase::latin());
$result = $recognizer->recognize(Image::fromPng(file_get_contents('subtitle.png')));

$result->text();                    // all lines, joined with "\n"
$result->confidence();              // mean confidence of all glyphs, from 0 to 1
$result->lines[0]->text;            // the first line
$result->lines[0]->confidence;      // mean confidence of the glyphs on the first line
$result->lines[0]->chars[0];        // RecognizedChar: text, confidence, italic, glyph, sample
  • Images: Image::fromPng($bytes) decodes every PNG colour type without GD. Image::fromGd($gdImage) copies a GD image. Image::fromRgba($pixels, $width, $height) takes raw RGBA bytes or a list of integers.
  • Confidence: the share of the database glyph's line points that agree with the image, from 0 to 1. An unknown glyph reads as * with confidence 0.
  • Ink: a pixel is ink when the sum of its premultiplied red, green and blue is at least inkThreshold (200). So white or yellow text counts and a black outline does not.
  • State: a recognizer learns the glyph heights from the images it reads and uses them to split the next images. Use one recognizer per subtitle stream, or call reset().
Option Default Meaning
inkThreshold 200 Minimum premultiplied red + green + blue of an ink pixel, 1 to 765
spaceWidth null Empty columns that make a space. null uses a third of the typical glyph height of the line
maxWrongPixels 25 Error budget of the loose match passes
fixLatinCase true Picks upper or lower case for letters such as o and O from their height
unknownText '*' Text of a glyph that matches nothing
italicSlant 0.0 Above 0, a glyph that matches nothing is slanted back by this factor and tried again
rightToLeft false Puts the glyphs of each line in right to left order
minLineHeight 12 Minimum line height in pixels until the recognizer has learned the glyph heights

Databases and training

GlyphDatabase reads and writes the .nocr files of Subtitle Edit, version 1 and 2. The package ships the Latin database of Subtitle Edit, 699 glyphs in 477 KB. Other scripts and other fonts need their own database.

Train a glyph from a sample that a person confirmed. The new glyph goes first, so it wins over older glyphs that match equally well:

use GlyphOcr\Trainer;

$database = GlyphDatabase::fromFile('my-font.nocr');
$trainer = new Trainer();
foreach ($recognizer->recognize($image)->unknownChars() as $char) {
    $database->add($trainer->train($char->sample, askAPerson($char->sample->toAscii())));
}
$database->save('my-font.nocr');
  • Recognizer::split($image) returns the glyph samples of each line without matching them, with null for a space.
  • GlyphSample::merge([$left, $right]) joins the parts of a character that the splitter cuts apart, such as " or %.
  • The trainer draws random line segments with a seeded generator, so the same sample and seed give the same glyph.

Accuracy

The golden images in tests/fixtures are white or yellow subtitles in DejaVu Sans and Liberation Sans, 20 to 60 pixels, 1 or 2 lines. tests/fixtures/SOURCES.md describes them. One recognizer reads each set in order.

Set Latin database: characters Latin database: lines After training: characters After training: lines
Blu-ray, smooth RGBA 89.7% 2 of 11 95.1% 4 of 11
PGS palette 90.2% 2 of 8 98.6% 6 of 8
DVD, 4 colours, 24 to 30 px 61.9% 0 of 8 91.0% 3 of 8
Italic 82.2% 0 of 5 95.0% 3 of 5
No outline 77.5% 1 of 5 93.3% 1 of 5
Small, 20 to 24 px 46.2% 0 of 4 94.2% 2 of 4
All 79.1% 5 of 41 94.6% 19 of 41
  • Character accuracy is 1 minus the edit distance divided by the length of the expected text.
  • The Latin database has no glyphs of these two fonts. The numbers show how it does on fonts it has not seen.
  • "After training" adds one trained glyph per character for each font, size and style: the alphabet, digits and .,!?'-:, drawn on a separate image. Most remaining errors are I and l, which have the same shape in sans-serif fonts.
  • With italicSlant: 0.2, the italic set reads 86.1% of the characters.
  • With the same database, the port reads every golden image exactly as Subtitle Edit does. SubtitleEditParityTest checks this.

Speed

Mean time per golden image, one image of 1 or 2 lines, after the database is loaded:

PHP 8.2 PHP 8.5
Latin database 280 ms 272 ms
After training 92 ms 76 ms
Load the Latin database 119 ms 54 ms

A full 1920x1080 frame with the same 2 lines takes about 1 second and 103 MB, so crop the image to the text where you can. A glyph that matches nothing is the slow case, because it runs through every match pass. Measured on one core of an x86_64 machine, without JIT.

Limits

  • One text colour on a transparent or dark background. The ink threshold removes the outline, so text with a light outline or a light background does not work.
  • The matcher compares shapes. Glyphs with the same shape in a font, such as I and l in most sans-serif fonts, come out as whichever the database has first.
  • No dictionary and no language model. Subtitle Edit fixes common OCR errors in a separate step, which this package does not port.
  • Italic text needs italic glyphs in the database, or italicSlant.

Exceptions

Every exception implements GlyphOcr\Exceptions\GlyphOcrException: InvalidArgumentException, InvalidImageException for a PNG that does not decode, and InvalidDatabaseException for a .nocr file that does not load.

Attribution

This package is a port of the nOCR engine from Subtitle Edit by Nikolaj Olsson, MIT license, at commit b8da12a. LICENSE keeps both copyright notices.

This package Subtitle Edit file in src/libuilogic/Ocr
Internal/Matcher.php, Glyph.php NOcrDb.cs, NOcrChar.cs
GlyphDatabase.php NOcrDb.cs, NOcrChar.cs (file format)
Internal/Splitter.php, Internal/SplitItem.php NikseBitmapImageSplitter2.cs, ImageSplitterItem2.cs
Internal/Bitmap.php NikseBitmap2.cs
Internal/LinePoints.php NOcrLine.cs
Internal/CaseFixer.php NOcrCaseFixer.cs
Internal/LineHeightTracker.php OcrLineHeightTracker.cs
Trainer.php NOcrChar.cs, NOcrLineGenerator.cs
GlyphSample.php (merge) ExpandedOcrGroup.cs
resources/Latin.nocr Ocr/Latin.nocr at the repository root, unchanged