Evaluation method
How to evaluate enterprise document AI
A credible enterprise AI evaluation starts before the demo: freeze a representative set of documents, write expected evidence and failure cases, then score each system layer separately.
The method
Five things to decide and write down
Change one thing at a time. Swap the documents, the model, the prompt or the permissions and it is a different test.
- 01
Define one workflow
Name the person, decision or output, document set and point at which human review remains required.
- 02
Freeze the test document set
Record the files, versions, permissions and expected evidence before running questions.
- 03
Write the question set
Include known-answer, absent-answer, conflicting-source and access-boundary cases alongside the polished question a demo would use.
- 04
Score the layers separately
Measure retrieval, citation support, answer support and reviewer usefulness independently so one strong layer cannot hide another.
- 05
Record the decision
Keep configuration, raw results, reviewer notes, limitations, owner and next action in a reproducible evaluation record.
Writing the questions
Make the difficult cases part of the brief
| Test class | Fixture | Expected behaviour |
|---|---|---|
| Known answer | A relevant current source exists. | Find the expected evidence and support the material claims. |
| Absent answer | The answer is not in the approved documents. | Decline, qualify or make the evidence gap visible. |
| Conflicting sources | Two sources disagree or differ by version. | Expose the conflict and the source/version context. |
| Plausible distractor | Related wording exists but does not answer the question. | Avoid treating topical similarity as claim support. |
| Access boundary | A source exists outside the test identity's approved scope. | Do not retrieve or reveal the restricted material. |
Do not hide a weak layer inside one score
Agree what counts as a pass before you see the results, and write down where reviewers disagreed.
- M01Retrieval relevance
- Did the candidate passages contain the evidence a reviewer expected?
- M02Citation correctness
- Does each cited passage support the claim attached to it?
- M03Citation completeness
- Are material answer claims supported, or are important claims uncited?
- M04Answer support
- Is the answer entailed by the selected evidence without unsupported additions?
- M05No-answer behaviour
- When the documents do not answer the question, or disagree, does it say so?
- M06Control behaviour
- Did permissions, document scope and the agent's tools behave as agreed?
- M07Reviewer usefulness
- Can the person responsible check and correct the answer without piecing the sources back together?
Decision record
Keep enough evidence to reproduce the conclusion
The record should name the documents and versions used, who ran it, how it was set up, which model, the date, the raw answers, the scores and who owns the decision.
When this method is not enough
- Clinical, legal, financial and safety work may need further validation by a qualified person, and regulatory sign-off.
- A small evaluation document set does not establish performance across every document type, language or permission model.
- A citation shows you where to look. The reviewer still checks that the source says what the answer claims, and what is missing.
- Whether a deployment suits you is a separate architecture and security review.
Plan an evaluation with us
Describe the workflow and your constraints in the form. Nothing confidential is needed to start; we agree which documents to use, and how to send them, afterwards.