Skip to content

Legal evaluation guide

AI for legal document work

How these systems read legal documents, which failures survive a skim, and how to test one on your own matter files before you commit to it.

Updated 21 Aug 2026

Definition

Legal document work covers research, drafting, review and analysis. These are different tasks with different failure modes, and a tool that is strong at one can be weak at another.

01

Separate the four tasks before comparing anything

"Legal AI" is sold as one category and bought as one purchase, but it describes at least four jobs. Drafting produces new wording. Analysis answers questions across a document estate you already hold.

The distinction decides what evidence matters. A research tool lives or dies on the breadth and currency of its source documents. An analysis tool lives or dies on whether it retrieves the right passage from your own material and shows you where it came from. Comparing them on a single feature grid produces a purchase nobody can defend.

  • Research finds authority outside your files
  • Drafting produces new wording
  • Review checks a document against a standard
  • Analysis answers questions across your own estate
02

Where these systems actually fail

The widely reported failure is invented authority: a citation to a case or clause that does not exist. It is the most visible failure and the easiest to test for, which makes it the least dangerous.

The failures that survive review are quieter. A system retrieves a superseded version of a clause and answers confidently from it. It answers from one document when two in the estate disagree, and never mentions the second. It returns nothing relevant and produces a fluent answer anyway. Each of these produces output that reads correct and passes a skim.

  • Invented authority, visible and testable
  • Superseded versions answered as current
  • Conflicting sources silently resolved
  • Fluent answers where evidence is absent
03

What a citation has to do to be worth anything

A citation is only useful if it can be checked at the level of the individual claim. A footnote pointing at a 200-page document is a gesture. A reference that opens the source, moves to the page and marks the passage is a control.

Test citations claim by claim rather than answer by answer. Split an answer into its material propositions, then ask of each one whether a cited passage supports it, whether the passage is from the document and version the citation names, and whether material claims exist with no citation at all. Correctness and completeness are separate measurements and should not be collapsed into one score.

  • Claim-level, not answer-level
  • Source, page and passage identifiable
  • Correct document version named
  • Uncited material claims counted separately
04

Documents that break naive systems

Legal estates are unusually hostile to document processing. Executed agreements arrive as scans. Obligations sit in tables where a single cell carries the meaning. Amendments change wording without replacing the document that contains it, so the current position is spread across a chain.

Test on the difficult material rather than a clean sample. Include a poor scan, a document whose substance is in a table, an agreement amended twice, a set with a superseded version still present, and a question whose answer is genuinely not in the document set.

  • Scanned and poor-quality originals
  • Obligations carried in table cells
  • Amendment chains and current position
  • Superseded versions still in the estate
  • Questions with no answer in the document set
05

Privilege and information barriers are an architecture question

A policy that says users should not see other matters is not a control. The question is whether the retrieval layer can return a passage from a document the user has no right to read, and whether a generated answer can carry that content even when the source is not shown.

Ask where the boundary is enforced. Enforcement at the point of retrieval behaves differently from filtering applied to results, and differently again from separate storage per matter. Test it directly: revoke a user's access to a source and ask a question whose answer only that source contains.

  • Enforcement point, not stated policy
  • Retrieval-time versus post-filter
  • Generated output can leak what it does not cite
  • Test by revoking access and re-asking
06

The positions a UK firm needs written down

Regulatory expectations here are about competence and supervision rather than a prohibition. A solicitor remains accountable for work produced with a tool, which means the firm needs a recorded position on who checks output, against what, and before which decisions.

Under UK GDPR the questions are the ordinary ones asked precisely: what personal data enters the system, which processors handle it, where it is held, how long it stays, and what happens on deletion. Answer them per deployment configuration, because the answers change with the configuration and a generic vendor statement will not survive a DPIA.

  • Named reviewer and review point
  • Competence and supervision recorded
  • Processors and locations per configuration
  • Retention and deletion answered specifically
07

Evaluate on your own documents before buying

Vendor demonstrations run on corpora chosen to succeed. The only evaluation that predicts anything is one on your own files, with questions written before you see any output and expected evidence recorded in advance.

Keep the reviewer's disposition per question rather than a headline accuracy figure. Counts are only meaningful alongside the denominator, the scoring rules and the cases that failed.

  • Your documents, not the vendor's sample
  • Questions and expected evidence written first
  • what the reviewer decided recorded per question
  • Failed cases published with the counts

What this page does not prove

  1. B1Marella supports research and document analysis. It does not provide legal advice or replace a qualified lawyer.
  2. B2Citation presence is not citation correctness.
  3. B3Answers stream before the fact-check completes, so the check flags rather than blocks.
  4. B4Evaluation results apply only to the document set and configuration tested.

Frequently asked questions

How does AI analyse a legal document?

A system splits documents into passages, indexes them for retrieval, selects passages it judges relevant to a question and supplies those to a model that produces an answer. Answer quality therefore depends on retrieval before it depends on the model: if the wrong passages are selected, a strong model produces a fluent answer from the wrong evidence.

How accurate is AI on legal questions?

There is no single accuracy figure that transfers between corpora. Accuracy depends on the documents, the questions asked and the configuration tested, so a number quoted without those three is not evidence. Measure citation correctness and citation completeness separately on your own material, and record the failures alongside the successes.

Can AI give legal advice?

No. These systems support research and document analysis. Interpretation, judgement and advice remain with a qualified professional, and the firm remains accountable for work produced with the tool.

How do you stop an AI system crossing an information barrier?

By enforcing access at the point of retrieval rather than filtering results afterwards, and by testing it: revoke a user's access to a document and ask a question whose answer only that document contains. Anything less is a stated policy rather than a control.

Test the claim on your documents

Pick a real piece of work, agree what a good answer looks like, then go through the results together.