Skip to content

Evaluation method

How to evaluate enterprise document AI

A credible enterprise AI evaluation starts before the demo: freeze a representative set of documents, write expected evidence and failure cases, then score each system layer separately.

The method

Five things to decide and write down

Change one thing at a time. Swap the documents, the model, the prompt or the permissions and it is a different test.

  1. 01

    Define one workflow

    Name the person, decision or output, document set and point at which human review remains required.

  2. 02

    Freeze the test document set

    Record the files, versions, permissions and expected evidence before running questions.

  3. 03

    Write the question set

    Include known-answer, absent-answer, conflicting-source and access-boundary cases alongside the polished question a demo would use.

  4. 04

    Score the layers separately

    Measure retrieval, citation support, answer support and reviewer usefulness independently so one strong layer cannot hide another.

  5. 05

    Record the decision

    Keep configuration, raw results, reviewer notes, limitations, owner and next action in a reproducible evaluation record.

Writing the questions

Make the difficult cases part of the brief

Do not hide a weak layer inside one score

Agree what counts as a pass before you see the results, and write down where reviewers disagreed.

M01Retrieval relevance
Did the candidate passages contain the evidence a reviewer expected?
M02Citation correctness
Does each cited passage support the claim attached to it?
M03Citation completeness
Are material answer claims supported, or are important claims uncited?
M04Answer support
Is the answer entailed by the selected evidence without unsupported additions?
M05No-answer behaviour
When the documents do not answer the question, or disagree, does it say so?
M06Control behaviour
Did permissions, document scope and the agent's tools behave as agreed?
M07Reviewer usefulness
Can the person responsible check and correct the answer without piecing the sources back together?

Decision record

Keep enough evidence to reproduce the conclusion

The record should name the documents and versions used, who ran it, how it was set up, which model, the date, the raw answers, the scores and who owns the decision.

When this method is not enough

  • Clinical, legal, financial and safety work may need further validation by a qualified person, and regulatory sign-off.
  • A small evaluation document set does not establish performance across every document type, language or permission model.
  • A citation shows you where to look. The reviewer still checks that the source says what the answer claims, and what is missing.
  • Whether a deployment suits you is a separate architecture and security review.

Plan an evaluation with us

Describe the workflow and your constraints in the form. Nothing confidential is needed to start; we agree which documents to use, and how to send them, afterwards.