Skip to content

Open evaluation protocol

Regulated Document AI Evaluation Kit

A free protocol and blank worksheets for testing cited answers on your own documents, including what a system does when the evidence is not there.

Definition

To evaluate citation correctness and completeness, split each answer into its material claims, check every citation against the source passage and version it names, then count the material claims left with no citation at all. Version 1.0 publishes that protocol and the blank instruments, with no product score and no vendor ranking.

01

Select a representative document set

Sample the document formats, versions, permissions, dates, quality and language conditions present in the intended workflow. Record exclusions and source rights.

  • Current controlling sources
  • Superseded or stale versions
  • Difficult scans, tables and appendices
  • Restricted and cross-department sources
  • Plausible distractors
  • Documents with known conflicts
02

Prepare reference questions

Write questions, expected evidence and acceptable failure behaviour before testing the system.

  • Known-answer lookup
  • Multi-source synthesis
  • Conflicting evidence
  • Stale-source trap
  • No-answer request
  • Permission-boundary case
03

Score citations claim by claim

Split the answer into material claims. For each citation, record source identity, version, passage access and whether it fully, partly or does not support the attached claim. Then mark material uncited claims separately.

  • Split into material claims
  • Record source identity and version
  • Full, partial or no support
  • Mark material uncited claims
04

Score abstention and uncertainty

When the evidence is not enough, record what the system did: asserted something it could not support, qualified the answer, asked a question back, showed the conflict, escalated, or declined. Also record incorrect refusals when evidence was adequate.

  • Unsupported assertion or qualification
  • Clarification or conflict display
  • Safe escalation or decline
  • Incorrect refusals recorded
05

Test permissions

Use approved test identities and non-sensitive fixtures. Verify allowed retrieval, denied retrieval, revoked access, source deletion and cross-organisation separation without leaking restricted content into notes or analytics.

  • Approved identities and fixtures
  • Allowed and denied retrieval
  • Revoked access and deletion
  • No leakage into notes
06

Review data flow and deployment

Assign an operator and evidence owner to every node.

  • Source, parsing and retrieval
  • Model, output and logs
  • Backups, retention and deletion
  • Operator and evidence owner
07

Use the evaluator worksheet

The downloadable CSV records question, reference evidence, identity scope, document set and configuration, retrieved sources, citation dimensions, answer support, uncertainty, permission behaviour and reviewer disposition.

  • Question and reference evidence
  • Identity scope and configuration
  • Retrieved sources and citations
  • Uncertainty and reviewer disposition
08

Use the procurement scorecard

Convert test findings and open architecture questions into evidenced, open, accepted-with-owner or blocking decisions. Do not average away a hard requirement.

  • Test findings and open questions
  • Evidenced or open
  • Accepted-with-owner or blocking
  • No averaging of hard requirements
09

Limitations and version history

Version 1.0 was published on 15 July 2026. It contains no benchmark results. Adaptation can change comparability; record every local scoring change, document set version, configuration, evaluator instruction and failed run.

  • Version 1.0, no benchmark results
  • Adaptation changes comparability
  • Record scoring and document set changes
  • Record configuration and failed runs

Version 1.0 · free to use

Six tests, one written decision

Prepare reference evidence before testing. Score the sources and the answer separately, and record what the reviewer decided.

  1. T01

    Known answer

    The expected answer and controlling passage are recorded before the run.

  2. T02

    Synthesis

    The answer needs several sources and material claims are scored separately.

  3. T03

    Conflict

    Approved sources disagree and the output should expose rather than flatten the conflict.

  4. T04

    Stale source

    A superseded source is present beside a current source with explicit dates.

  5. T05

    No answer

    The document set lacks adequate support and acceptable decline or qualification is defined.

  6. T06

    Permission

    Relevant evidence exists outside the test identity's approved scope.

Open worksheets · CSV

Readable in a spreadsheet or plain-text editor; no gated download and no macros.

What this page does not prove

  1. B1The kit is a method, not certification or regulatory approval.
  2. B2Blank worksheets contain no Marella AI performance result.
  3. B3Evaluation must respect document licensing, confidentiality and access rules.
  4. B4A small document set does not establish universal performance.

Frequently asked questions

Can the kit be used for another vendor or an internal build?

Yes. Use the same documents, question set, scoring rules and reviewers for each system, then record configuration differences.

Does the kit produce one overall score?

No default aggregate is prescribed. Keep hard constraints and citation, answer, uncertainty, permission and operating dimensions visible rather than hiding them in one number.

Why are there no benchmark results?

A defensible result needs a publishable documents, reference answers, configuration, evaluator process, failed runs and limitations. This release publishes the protocol without inventing data.

Run the same protocol on Marella AI

Agree the representative workflow, expected evidence and safe corpus-handling route before testing.