---
title: "What is multimodal RAG?"
description: "Multimodal RAG retrieves pages, frames, or transcripts — not only plain-text passages — and passes those pieces to the model."
url: "https://pixeltable.com/learn/what-is-multimodal-rag"
updated: "2026-09-29"
vertical: "Retrieval"
doc: "https://docs.pixeltable.com/datastore/embedding-index"
---

# What is multimodal RAG?

Multimodal RAG retrieves pages, frames, or transcripts — not only plain-text passages — and passes those pieces to the model.

Updated: 2026-09-29
Part of [What is RAG?](https://pixeltable.com/learn/what-is-rag).

## On this page

- [How it works](https://pixeltable.com/learn/what-is-multimodal-rag#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-multimodal-rag#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-multimodal-rag#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-multimodal-rag#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-multimodal-rag#questions)

## How it works {#how-it-works}


- Each modality is stored as a typed value.
- Retrieval units are passages, frames, or transcript windows.
- The model sees the retrieved piece and a pointer to its source.

## What it is not {#what-it-is-not}

It is not text-only RAG with images stored as unused attachments.

## multimodal RAG: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Evidence | Page, frame, or transcript | Text chunks only |
| Unused media | A failure of the pipeline | Normal, if images are just attachments |
| Index | One per modality you intend to search | A single text index over filenames |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable keeps each modality on the table and retrieves from the view that holds the pieces. The model call is another column or a query.

## Questions {#questions}

### How does multimodal RAG work? {#faq-1}

Each modality is stored as a typed value. Retrieval units are passages, frames, or transcript windows. The model sees the retrieved piece and a pointer to its source.

### What is multimodal RAG often confused with? {#faq-2}

It is not text-only RAG with images stored as unused attachments.

## In the blog

- [Powering Multimodal LLMs: Pixeltable MCP Servers for Rich Data Context](https://pixeltable.com/blog/pixeltable-mcp-servers)
- [VideoRAG: Where Pixeltable Stores Frames and Indexes](https://pixeltable.com/blog/videorag-where-pixeltable-stores-indexes)

## Related

- [Documentation](https://docs.pixeltable.com/datastore/embedding-index)
- [Production Multimodal RAG with Pixeltable: Dev to Deployment](https://pixeltable.com/blog/multimodal-rag-production)
- [Production RAG](https://pixeltable.com/use-cases/production-rag-implementation)
- [Multimodal application](https://pixeltable.com/use-cases/multimodal-ai-application-development)
