Skip to content

Glossary

What is data extraction?

Data extraction finds chosen values, entities or relationships in source material and turns them into structured output another system can use.

Updated 21 Aug 2026

01

Where this one gets misread

Extraction demos run on clean examples. The figure that matters is accuracy on your worst documents, not the average across a curated set. Ask about the scans, the amendments and the forms someone filled in by hand.

02

Questions to ask

Ask for a worked example on your own material, and the evidence needed to reproduce it.

  • Which fields, and what happens when one is absent?
  • How are low-confidence extractions surfaced rather than guessed?
  • Who corrects an error, and does the correction persist?
  • What was accuracy measured on?
03

How Marella uses the term

We use “Data extraction” only where a product mechanism or an evaluation method backs it up, and we say when the behaviour depends on how a deployment is configured.

  • Backed by a product mechanism or an evaluation method
  • Deployment differences flagged

What this page does not prove

  1. B1A definition is not a claim about how the product performs.
  2. B2Vendor implementations vary.
  3. B3Test the term against a representative workflow.

Test the claim on your documents

Pick a real piece of work, agree what a good answer looks like, then go through the results together.