Ask a general-purpose chatbot about your retention policy and you’ll get something fluent and well organised back. None of that tells you the answer is right, or current, or taken from the policy you meant. A cited answer beats a confident one because you can check it.
Where the answer feeds work someone will be held to, the reviewer needs a route back to the source that takes seconds rather than an afternoon. Anything they can’t check shouldn’t be treated as settled.
What a cited answer changes
A citation earns its place by pointing at the passage that’s supposed to support the claim, rather than at the document the passage sits inside. Where the parsed format holds reliable coordinates, the product can open that location beside the answer and the reviewer reads the sentence instead of hunting for it. That saves the hunting. It doesn’t save the judgement.
Cheap checking matters because a check nobody performs isn’t a control at all. A reviewer who can open the passage in one click will catch a citation that doesn’t support the claim it’s attached to, or a source that was superseded two years ago and never withdrawn.
It also makes the behaviour on unanswerable questions something you can test. A system can be set up to decline or hedge when it hasn’t found much, and the only way to find out whether yours does is to ask it things the document set genuinely doesn’t cover. When someone later asks how you knew, the citations are what you show them. They’re part of an audit record, not all of it.
What to test in practice
Write the questions down before the demo rather than during it. You want some with known answers, some that need two or three passages pulled together, some the document set genuinely can’t answer, and some where you already know two sources contradict each other. For every answer, record whether the cited passage supports the claim, whether the claims that matter carry a citation at all, and whether the source is still in force.
Marella AI shows the citation and lets you open the source behind it. How well that works depends on the file, on what the search returned and on how your deployment is set up. Run the evaluation on the document set you would actually deploy against.
The buying test
If a wrong answer costs you something, put this in front of the vendor early: ask a question out of your real work, then open the passage behind each claim that matters.
Score the correctness of the citation, how much of the answer carries one, and what happens when nothing in the document set supports the question, and keep those three scores apart. A citation can look perfect on screen and still point at weak evidence, so a single overall score will hide the thing you most need to see.