Retrieval

What is multimodal RAG?

Multimodal RAG retrieves pages, frames, or transcripts — not only plain-text passages — and passes those pieces to the model.

Updated · Part of What is RAG?

How it works

  • Each modality is stored as a typed value.
  • Retrieval units are passages, frames, or transcript windows.
  • The model sees the retrieved piece and a pointer to its source.

What it is not

It is not text-only RAG with images stored as unused attachments.

multimodal RAG: this, and the thing it is confused with

multimodal RAG: this, and the thing it is confused with
ThisNot this
EvidencePage, frame, or transcriptText chunks only
Unused mediaA failure of the pipelineNormal, if images are just attachments
IndexOne per modality you intend to searchA single text index over filenames

Where Pixeltable fits

Pixeltable keeps each modality on the table and retrieves from the view that holds the pieces. The model call is another column or a query.

Questions

How does multimodal RAG work?
Each modality is stored as a typed value. Retrieval units are passages, frames, or transcript windows. The model sees the retrieved piece and a pointer to its source.
What is multimodal RAG often confused with?
It is not text-only RAG with images stored as unused attachments.