Skip to content

Blog

How accurate is AI on legal questions?

An accuracy figure means little without the corpus and the questions behind it. What to measure instead, and why the no-answer cases tell you most.

Reuben McQueenlegalbuying-ai
A page of marked-up text beside a notebook.

“How accurate is it?” is the first thing everyone asks, and it’s the question I find hardest to answer honestly, because any single figure quoted without its documents, its questions and its configuration doesn’t actually tell you anything.

Accuracy isn’t a property a system carries around with it. It’s what happened when one system ran against one set of documents, answering one set of questions, in one configuration. Swap the document set and the number moves. Ask different questions and it moves again. When a vendor gives you a benchmark, they’re telling you how their system did on documents they chose, which is the condition it looks best under.

That isn’t a reason to stop measuring. Measure the things that transfer, on your own material.

It’s four measurements, not one

The trouble with a single accuracy score is that it hides the failures you’d most want to see.

The first is citation correctness: for every citation, does the passage it points at genuinely support the claim attached to it? And check the document identity while you’re there, because a passage can support a claim beautifully and still have come from a version that was superseded eighteen months ago.

The second is citation completeness, which is a different question entirely. Are there claims in the answer with no citation at all? You can have an answer where every single citation checks out and the bulk of the text is still unsupported assertion sitting between them.

Third is what happens when there’s nothing to find. A system that produces a confident, fluent answer out of an empty document set is worse than one that throws an error, because at least the error is visible.

And fourth, what happens when two documents disagree. Does the answer surface the disagreement, or quietly pick one and move on? Real document estates are full of things that contradict each other. That’s the ordinary case.

A system can do well on the first of those and badly on the third, and the average hides both.

Score claims, not answers

Break the answer into the statements a decision would actually rest on, and score each one separately.

The reason is that support is almost never uniform across a paragraph. Two sentences are well cited, the third is an inference the sources don’t quite carry, the fourth is background the model supplied from its training data. Give the paragraph one score and all of that disappears, and it’s the fourth sentence that eventually causes you a problem.

This is slower than reading answers and forming an impression, which is what most evaluations actually are. It’s also the only version that produces a number you’d be willing to defend to a regulator.

Write the questions first

Once you’ve watched a system answer, you can’t unsee it, and question design starts drifting towards the things it handles well. Nobody does it deliberately; it just happens.

So write the question set before you run anything, and write down the evidence you expect each answer to rest on. Get a spread in there: questions with one clear answer, questions that need synthesis across several documents, questions whose answer is buried in a table or a footnote, questions where the current position is spread across an amendment chain, at least one the document set can’t answer, and at least one where two sources conflict.

Publish the denominator

A rate without its denominator is decoration. “94% citation correctness” starts to mean something when it arrives with the number of claims scored, the scoring rules, who did the scoring and which cases failed. Without those it’s a marketing number and should be read as one, including when we’re the ones quoting it.

A vendor who shows you the questions their system got wrong is handing you far more usable information than one who shows you a higher number, and it’s worth telling them so.

Where this leaves you

Run the evaluation yourself, on a bounded set of your own documents, before you buy rather than after. Keep the reviewer’s disposition question by question instead of collapsing to a headline. Expect an uneven result. Most systems do well on straightforward look-ups, less well when an answer has to be pulled from several documents, and the cases where nothing was there to find will tell you the most.

Most systems handle the easy questions perfectly well. What separates them is what they do when the evidence is missing or the sources disagree, and those are the cases that get the least attention.

The full method is in the guide to evaluating AI citations, with blank worksheets in the evaluation kit. If you are still working out which kind of system you are testing, legal AI tools compared sets out how research, drafting and review tools differ, and which AI is best for legal research covers the research tools specifically.

Reuben McQueen

Written by Reuben McQueen, Co-founder & CTO at Marella AI.

Share this article