Friday, 9 October 2026 SourcesAbout🌓
🇬🇧 UK ▾
BREAKING
› Man Utd held back as Spurs spent big - Different strategies collide on Saturday› Nuno on West Ham's bright start and being back in the Championship› Hunter Bell celebrates in Team GB's 'glam' female track success› Billam-Smith on Clarke revenge mission: 'I've got a lot of bitterness!'› Chelsea latest: Caicedo features in friendly as midfielder steps up recovery› Swiss Darts Trophy 2026: Schedule, draw, dates as Bunting defends his title› 'It's about time!' - F1 drivers excited amid Rwanda GP rumours› 'Stick together, enjoy the ride and smile' - Haaland's message to Man City fans› Campbell would end retirement to fight Benn: 'He was insulting me!'› Russell, Antonelli to race with different specs amid Mercedes upgrade concern› Man Utd held back as Spurs spent big - Different strategies collide on Saturday› Nuno on West Ham's bright start and being back in the Championship› Hunter Bell celebrates in Team GB's 'glam' female track success› Billam-Smith on Clarke revenge mission: 'I've got a lot of bitterness!'› Chelsea latest: Caicedo features in friendly as midfielder steps up recovery› Swiss Darts Trophy 2026: Schedule, draw, dates as Bunting defends his title› 'It's about time!' - F1 drivers excited amid Rwanda GP rumours› 'Stick together, enjoy the ride and smile' - Haaland's message to Man City fans› Campbell would end retirement to fight Benn: 'He was insulting me!'› Russell, Antonelli to race with different specs amid Mercedes upgrade concern
Technology

ABBYY gives old-school OCR a job in the AI pipeline

The Register ·
ABBYY gives old-school OCR a job in the AI pipeline

ABBYY has packaged its FineReader OCR engine as a self-hosted tool for turning documents into structured text that AI systems can use.

The company's new FineParser runs in a Docker container on a CPU, without requiring a GPU.

It aims to preserve the layout of a document as it extracts its contents.

It takes images of documents, in multiple languages, and turns them into structured, formatted text – so that they can be processed using, for instance, modern generative AI LLMs.

This isn't its sole purpose: the company suggested it could help bring print text into a modern CMS, or for archiving, or as a stage in some form of production pipeline.

ABBYY describes FineParser's approach as "deterministic" AI.

It extracts text and document structure rather than generating a plausible rendition of them.

Its output can then be passed to a generative AI system, whose responses are less predictable.

The company also offers its own programmable machine-learning framework, NeoML, which is FOSS and available on GitHub.

It also publishes an OCR SDK for companies wanting to embed the FineReader engine into their own products.

FineParser itself isn't open source, although ABBYY maintains a GitHub repository for examples and community support.

The self-hosted tool has a free tier allowing 1,000 pages per month for one year.

ABBYY says its subscription tiers connect to a license server for validation; a fully offline deployment requires an Enterprise plan.

Preserving structure means recognizing columns in reading order, headings, and tables – even those without borders – rather than returning a jumble of extracted text.

Read the full article on The Register ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.

More from The Register

See all ›

More in Technology

See all ›