Skip to content

Glossary

What is latency?

Latency is the time between a request and an agreed response point. A useful figure says what was measured and gives percentiles, not just an average.

Updated 21 Aug 2026

01

Where this one gets misread

A latency figure without percentiles and a definition of what was timed is decoration. Time to first token and time to a complete, checked answer are very different numbers, and the flattering one is usually the one quoted.

02

Questions to ask

Ask for a worked example on your own material, and the evidence needed to reproduce it.

  • What exactly was timed, from which event to which?
  • What are p50 and p95, on what corpus size?
  • Under how many concurrent users?
  • Does any citation or verification step finish inside that figure?
03

How Marella uses the term

We use “Latency” only where a product mechanism or an evaluation method backs it up, and we say when the behaviour depends on how a deployment is configured.

  • Backed by a product mechanism or an evaluation method
  • Deployment differences flagged

What this page does not prove

  1. B1A definition is not a claim about how the product performs.
  2. B2Vendor implementations vary.
  3. B3Test the term against a representative workflow.

Test the claim on your documents

Pick a real piece of work, agree what a good answer looks like, then go through the results together.