Skip to content

Glossary

What is optical character recognition (OCR)?

Optical character recognition turns text in an image or scan into machine-readable characters. Whatever it gets wrong is carried into every later step.

Updated 21 Aug 2026

01

Where this one gets misread

OCR errors do not announce themselves. A misread figure or date flows into an answer and a citation that both look entirely correct, and the only way to catch it is to open the page.

02

Questions to ask

Ask for a worked example on your own material, and the evidence needed to reproduce it.

  • What is the character error rate on your worst scans?
  • How are low-confidence regions marked?
  • Is the original page viewable beside the extracted text?
  • What happens with handwriting, stamps and marginalia?
03

How Marella uses the term

We use “Optical character recognition (OCR)” only where a product mechanism or an evaluation method backs it up, and we say when the behaviour depends on how a deployment is configured.

  • Backed by a product mechanism or an evaluation method
  • Deployment differences flagged

What this page does not prove

  1. B1A definition is not a claim about how the product performs.
  2. B2Vendor implementations vary.
  3. B3Test the term against a representative workflow.

Test the claim on your documents

Pick a real piece of work, agree what a good answer looks like, then go through the results together.