Help centre
Preparing documents for ingestion into Marella
What to load first, what to put through the pipeline, how to spot the files that will not read cleanly, and how to keep the set in good shape.
What to load first
Start with the documents people actually ask about: the current policies, procedures, playbooks and precedents. A small set of live documents beats a huge archive of might-be-relevant ones, and you can always widen the scope once the first team trusts the answers.
What to test with the pipeline
Real corpora tend to contain:
- Scanned PDFs and image-heavy files, including the poor-quality scans
- Long documents, appendices and mixed formats
- Spreadsheets and structured exports
- Legacy files that predate whoever now owns them
Supported formats and OCR quality vary with file structure and configuration. Reconcile the source set against the media library, then look into anything missing, failed or only partly parsed.
Keeping the document set clean
Scope live policy sets rather than archives wherever you can, so answers come from what is actually in force. Connectors ingest what you point them at, so tidy source folders make the scoping precise. And where knowledge extraction is enabled, sample the entities and concepts against the source and agree how often that curation happens.
When in doubt
Put the awkward cases in the evaluation document set and write down what you expect them to do. Share the failures with your Marella AI contact so the scope gets looked at properly.