Multimodal data
What is a multimodal database?
A multimodal database stores structured values and typed media in one schema. An image, a video, an audio clip, or a document is a column the engine understands, not a file path hidden in a string.
Updated
What it is
A row can hold a title, a number, and a video at the same time. The catalog records that the video column is video. Bytes can live in a media store or in object storage you already run. Queries filter the structured fields and read the media the row points at.
Transforms belong on that schema too. Extracting frames, transcribing audio, or embedding a passage are columns that run when the row changes. An index on those outputs stays attached to the same row, so search does not depend on a second copy you update by hand.
How it works
Insert or update a row. The engine stores the structured values, records a typed pointer to the media, and runs the declared transforms for that row. Unchanged rows stay as they are. A later query can filter, play the media, or search an embedding that was computed from it.
- The schema names the media type, so frame extraction or document splitting is a property of the column, not a script that hopes the path is a video.
- Computed columns cache their outputs. A new file recomputes that row. Editing a prompt recomputes the columns that depend on it.
- History can sit on the same table, so you can see what the row looked like before a model or a file changed.
What it is not
The phrase is easy to confuse with four other things. None of them let you insert a clip and keep the derived frames, transcripts, and embeddings consistent.
- An academic benchmark. A “multimodal table” in a paper is often a dataset used to score models. It is not a system you insert into.
- A path in a string column. A URL is not media. Filtering the path does not give you frames, pages, or a duration.
- An untyped blob. If the engine does not know the value is video, you write the decoder yourself.
- A split stack. Files in a bucket, metadata in a warehouse, embeddings in another database, and a job to keep them aligned. That is several systems, not one schema.
Where the media lives, and what the engine can do with it
| Path or blob | Split stack | One schema | |
|---|---|---|---|
| Media in the row | Opaque path or untyped file | Bytes in a bucket, ids in SQL | A typed image, video, audio, or document column |
| Derived data | A script you remember to rerun | A pipeline between services | Columns that run when the row changes |
| Search | Filename or a separate index you load | Embeddings copied into another store | An index on a column of the same table |
| After an edit | Rerun the script | Reconcile every store | Recompute the rows that changed |
Where Pixeltable fits
Pixeltable is one multimodal database: Image, Video, Audio, and Document columns in a Python schema, with computed columns and embedding indexes on that schema. A longer essay on the table itself, as distinct from a benchmark or a file path, is What Is a Multimodal Data Table?
import pixeltable as pxtTableModel = pxt.model_base()class Library(TableModel, name='library'):title: pxt.Stringvideo: pxt.Videonotes: pxt.String | None
Questions
- Is a multimodal database the same as a vector database?
- No. A vector database stores embeddings and answers similarity queries. A multimodal database stores the media and the structured fields, and can keep an embedding index on those columns. The index is one part of the schema, not a second database you have to load.
- Where do the video and image bytes live?
- The catalog holds a typed pointer. The bytes live in the media store or in object storage. The column type is what lets later columns extract frames, transcribe audio, or split a document.
- What is a multimodal data table?
- A multimodal data table is the row-and-column form of this idea: structured values plus typed media, with computed columns and history on that schema. The essay is at /blog/what-is-a-multimodal-data-table.