Skip to content

Glossary

What is multimodal AI?

Multimodal AI takes in or produces more than one kind of content, such as text, images or audio. Which combinations work varies by model and workflow.

Updated 21 Aug 2026

01

Where this one gets misread

Multimodal gets stated as a property of the product when it is a property of one model on one step. Which combinations actually work varies by workflow, and the diagram in a deck rarely says which.

02

Questions to ask

Ask for a worked example on your own material, and the evidence needed to reproduce it.

  • Which inputs, on which step?
  • What happens to an image the model cannot read?
  • Is visual content cited the same way as text?
  • What was tested, on what material?
03

How Marella uses the term

We use “Multimodal AI” only where a product mechanism or an evaluation method backs it up, and we say when the behaviour depends on how a deployment is configured.

  • Backed by a product mechanism or an evaluation method
  • Deployment differences flagged

What this page does not prove

  1. B1A definition is not a claim about how the product performs.
  2. B2Vendor implementations vary.
  3. B3Test the term against a representative workflow.

Test the claim on your documents

Pick a real piece of work, agree what a good answer looks like, then go through the results together.