Skip to content

Research method

Enterprise AI search evaluation framework 2026

A transparent method for comparing retrieval, citations, unsupported-answer behaviour, permissions and operational fit on a representative corpus.

Updated 21 Aug 2026

Definition

This page publishes the method. It does not publish vendor rankings or unsupported market statistics.

01

Document set design

Use documents that reflect production formats, versions, permissions and quality.

  • Known-answer sources
  • Plausible distractors
  • Archived versions
  • Absent and conflicting evidence
02

Question set

Write expected evidence and acceptable no-answer behaviour before testing a system.

  • Lookup
  • Synthesis
  • Comparison
  • Unsupported request
03

Scoring

  • Retrieval relevance
  • Citation correctness and completeness
  • Answer support and abstention
  • Latency and reviewer usefulness
04

Reproducibility

  • Document set version and configuration
  • Models, providers and prompts
  • Date and reviewer guidance
  • Raw anonymised results

What this page does not prove

  1. B1We claim no benchmark result without a public documents, configuration and scoring record.
  2. B2Results on one document set do not establish universal performance.

Test the claim on your documents

Pick a real piece of work, agree what a good answer looks like, then go through the results together.