---
title: "Learn"
description: "Definitions for multimodal tables, RAG, video search, embeddings, agents, and cost."
url: "https://pixeltable.com/learn"
---

# Learn

Definitions first. Each blog tag maps to one of these pages, or to a release note or event that stays on the blog.

## Multimodal data

- [What is a multimodal database?](https://pixeltable.com/learn/what-is-a-multimodal-database) — A multimodal database stores structured values and typed media in one schema. An image, a video, an audio clip, or a document is a column the engine understands, not a file path hidden in a string.
- [What is a multimodal table?](https://pixeltable.com/learn/what-is-a-multimodal-table) — A multimodal table is a row-and-column object that holds structured values and typed media together, with derived columns and history on that same object.
- [What is a typed media column?](https://pixeltable.com/learn/what-is-a-typed-media-column) — A typed media column stores an image, video, audio clip, or document as a value the engine can decode, iterate, and index, instead of an opaque blob.
- [What is an image column?](https://pixeltable.com/learn/what-is-an-image-column) — An image column stores a still picture as a first-class value a schema can caption, embed, or transform.
- [What is a video column?](https://pixeltable.com/learn/what-is-a-video-column) — A video column stores a clip with a timeline, so frames, audio, and timestamps stay addressable from the row.
- [What is an audio column?](https://pixeltable.com/learn/what-is-an-audio-column) — An audio column stores a waveform or audio container the schema can transcribe, embed, or synthesize from.
- [What is a document column?](https://pixeltable.com/learn/what-is-a-document-column) — A document column stores a PDF or similar file so pages and passages can be split, embedded, and cited from the row.
- [What is a TableModel?](https://pixeltable.com/learn/what-is-a-tablemodel) — A TableModel is a Python class that declares tables, computed columns, and indexes. HTTP routes are a FastAPIRouter next to the class, not fields on it.
- [What is a view?](https://pixeltable.com/learn/what-is-a-view) — A view is a derived table whose rows are produced from a base table, often one row per frame, chunk, or other iterator output, and stay aligned when the base changes.
- [What is an iterator?](https://pixeltable.com/learn/what-is-an-iterator) — An iterator expands one source row into many child rows — frames, sentences, pages — and keeps a pointer back to the source.
- [What is a multimodal data plane?](https://pixeltable.com/learn/what-is-a-multimodal-data-plane) — A multimodal data plane is the layer that stores media, runs transforms, and serves retrieval as one system of record, rather than glue between a bucket, a warehouse, and an index.
- [What is an application database?](https://pixeltable.com/learn/what-is-an-application-database) — An application database stores the scalar rows of an app — users, orders, documents as JSON — and serves CRUD over them.

## Retrieval

- [What is RAG?](https://pixeltable.com/learn/what-is-rag) — Retrieval-augmented generation answers from a corpus you supply. The model sees retrieved passages, images, or other snippets at question time, not only what it learned in training.
- [What is an embedding?](https://pixeltable.com/learn/what-is-an-embedding) — An embedding is a numeric vector that places a passage, frame, or other piece in a space where similar items sit near each other.
- [What is an embedding index?](https://pixeltable.com/learn/what-is-an-embedding-index) — An embedding index ranks nearest neighbors for a query and stays attached to the table that owns those rows.
- [How does an embedding index differ from a vector database?](https://pixeltable.com/learn/embedding-index-vs-vector-database) — A vector database stores vectors and answers similarity queries. An embedding index is that search capability declared on the same table that holds the source pieces.
- [What is chunking?](https://pixeltable.com/learn/what-is-chunking) — Chunking splits a source into retrieval units — sentences, pages, token windows — small enough to embed and cite.
- [What is semantic search?](https://pixeltable.com/learn/what-is-semantic-search) — Semantic search ranks items by meaning, using embeddings, not only by exact keyword match.
- [What is cross-modal search?](https://pixeltable.com/learn/what-is-cross-modal-search) — Cross-modal search queries one modality with another — a sentence against frames, a still against a video library — because they share an embedding space.
- [What is multimodal RAG?](https://pixeltable.com/learn/what-is-multimodal-rag) — Multimodal RAG retrieves pages, frames, or transcripts — not only plain-text passages — and passes those pieces to the model.
- [What is a retrieval UDF?](https://pixeltable.com/learn/what-is-a-retrieval-udf) — A retrieval UDF packages a similarity query so an agent or another column can call retrieval as a function.

## Video

- [How does video search work?](https://pixeltable.com/learn/how-video-search-works) — Video search finds a moment inside a clip, not only a filename. The usual path is to sample frames, embed each frame, and rank those frames by similarity to a text phrase or a still image.
- [What is video RAG?](https://pixeltable.com/learn/what-is-video-rag) — Video RAG retrieves a moment — a frame, a timestamp, and maybe a transcript line — and gives that evidence to the model.
- [What is a frame iterator?](https://pixeltable.com/learn/what-is-a-frame-iterator) — A frame iterator walks a video at a chosen rate and emits one row per frame, with an index and a timestamp on the source clip.
- [What is keyframe extraction?](https://pixeltable.com/learn/what-is-keyframe-extraction) — Keyframe extraction selects representative frames — fixed rate or scene changes — so search and models run on stills instead of every encoded frame.
- [What is video intelligence?](https://pixeltable.com/learn/what-is-video-intelligence) — Video intelligence is a pipeline that turns clips into structured outputs you can search: frames, detections, transcripts, and embeddings.
- [What is a video moment?](https://pixeltable.com/learn/what-is-a-video-moment) — A video moment is a search hit that names the source clip plus a time offset, and often a frame, so you can jump to the match.

## Cost

- [What does an open-source multimodal AI stack cost?](https://pixeltable.com/learn/open-source-multimodal-ai-stack) — The bill is the engine, the hosted database if you use one, the media bytes, and the model APIs those pipelines call. A free license does not make the model calls free, and a low subscription does not make you a content-delivery network.
- [What is egress?](https://pixeltable.com/learn/what-is-egress) — Egress is network transfer of stored bytes out of a provider, usually billed per gigabyte.
- [How does incremental computation change AI cost?](https://pixeltable.com/learn/what-is-incremental-compute-cost) — Incremental compute cost means you pay model APIs mainly for new or changed rows. Cached columns are not sent again.
- [What is model API cost?](https://pixeltable.com/learn/what-is-model-api-cost) — Model API cost is the provider’s invoice for embeddings, transcription, vision, and chat. It is usually the surprise next to hosting.

## Documents

- [What is PDF RAG?](https://pixeltable.com/learn/what-is-pdf-rag) — PDF RAG keeps the file, splits it into passages or pages, embeds those units, and generates an answer that can cite them.
- [What is a document splitter?](https://pixeltable.com/learn/what-is-a-document-splitter) — A document splitter cuts a document into retrieval units — sentences, paragraphs, or pages — as rows over the source file.
- [What is OCR?](https://pixeltable.com/learn/what-is-ocr) — OCR, optical character recognition, turns pixels of text — scans, screenshots, photos — into strings.
- [What is document extraction?](https://pixeltable.com/learn/what-is-document-extraction) — Document extraction pulls text, layout, or tables out of a file so later columns can chunk, embed, or summarize.
- [What is document Q&A?](https://pixeltable.com/learn/what-is-document-qa) — Document Q&A answers questions from stored documents using retrieved passages, not a model that only saw the file once in its context window.

## Audio

- [What is speech-to-text?](https://pixeltable.com/learn/what-is-speech-to-text) — Speech-to-text transcribes spoken audio into text that can be searched, chunked, or passed to a model.
- [What is Whisper?](https://pixeltable.com/learn/what-is-whisper) — Whisper is a speech-recognition model family used to transcribe audio and video soundtracks into text.
- [What is text-to-speech?](https://pixeltable.com/learn/what-is-text-to-speech) — Text-to-speech synthesizes spoken audio from text.
- [What is audio intelligence?](https://pixeltable.com/learn/what-is-audio-intelligence) — Audio intelligence turns calls, podcasts, or voiceovers into chapters, search, and structured fields on top of a transcript.
- [What is an audio embedding?](https://pixeltable.com/learn/what-is-an-audio-embedding) — An audio embedding is a vector for a clip or a transcript window, used to find similar speech or sound rather than only exact words.

## Vision

- [What is computer vision?](https://pixeltable.com/learn/what-is-computer-vision) — Computer vision is the set of models and pipelines that detect, classify, segment, or search visual content.
- [What is CLIP?](https://pixeltable.com/learn/what-is-clip) — CLIP is a model that maps images and text into one space, so a sentence can retrieve a matching picture or frame.
- [What is object detection?](https://pixeltable.com/learn/what-is-object-detection) — Object detection finds instances of objects in an image or frame and returns boxes, usually with a class and a score.
- [What is YOLO?](https://pixeltable.com/learn/what-is-yolo) — YOLO is a family of real-time object detectors, including YOLOX, that predict boxes and classes on stills or video frames.
- [What is image segmentation?](https://pixeltable.com/learn/what-is-image-segmentation) — Image segmentation labels the pixels or regions that belong to an object or a prompt, rather than only a bounding box.
- [What is visual search?](https://pixeltable.com/learn/what-is-visual-search) — Visual search finds similar images or frames from a picture or a text description, using visual embeddings.
- [What is image annotation?](https://pixeltable.com/learn/what-is-image-annotation) — Image annotation is human or model labels — boxes, captions, classes — stored on visual rows so training and review can reuse them.
- [What is image generation?](https://pixeltable.com/learn/what-is-image-generation) — Image generation synthesizes or edits a picture from a prompt or another picture, and stores the result as an image.

## Agents

- [What is agent memory?](https://pixeltable.com/learn/what-is-agent-memory) — Agent memory is durable, queryable state an agent reads and writes across turns: messages, retrieved facts, and tool results, not a hidden JSON file.
- [What is MCP?](https://pixeltable.com/learn/what-is-mcp) — MCP, the Model Context Protocol, is a standard way for models and editors to call tools and read resources from a server.
- [What is a stateful agent?](https://pixeltable.com/learn/what-is-a-stateful-agent) — A stateful agent chooses its next action from persisted history and artifacts, not only from the current prompt.
- [What is an agent harness?](https://pixeltable.com/learn/what-is-an-agent-harness) — An agent harness is the data and tool layer an agent runs against — tables, memory, retrieval, and HTTP — so the loop is not just a prompt file.
- [What is an agent session log?](https://pixeltable.com/learn/what-is-an-agent-session-log) — An agent session log is a table of turns, tool calls, and outcomes you can query later.
- [What is a context graph?](https://pixeltable.com/learn/what-is-a-context-graph) — A context graph links what an agent retrieved, chose, and produced, so you can inspect why it acted.

## Orchestration

- [What is a declarative pipeline?](https://pixeltable.com/learn/what-is-a-declarative-pipeline) — A declarative pipeline is a schema of transforms. The engine decides when rows run, instead of you scheduling every step.
- [What is a computed column?](https://pixeltable.com/learn/what-is-a-computed-column) — A computed column is a value produced by an expression or a model when the row or its dependencies change, then cached.
- [What is a UDF?](https://pixeltable.com/learn/what-is-a-udf) — A UDF, a user-defined function, is your own logic the engine can call from a computed column or a query.
- [What is a UDA?](https://pixeltable.com/learn/what-is-a-uda) — A UDA, a user-defined aggregate, reduces many rows — for example the frames of one clip — into a summary with logic you write.
- [What is incremental computation?](https://pixeltable.com/learn/what-is-incremental-computation) — Incremental computation recomputes only the rows and downstream columns a change actually affects, instead of rerunning the whole pipeline.
- [What is a dependency graph?](https://pixeltable.com/learn/what-is-a-dependency-graph) — A dependency graph records which columns depend on which, so an edit invalidates only descendants.
- [What is an AI function?](https://pixeltable.com/learn/what-is-an-ai-function) — An AI function is a model call expressed in the schema — caption, transcribe, generate — rather than a separate pipeline product.
- [How does rate limiting work for model APIs?](https://pixeltable.com/learn/what-is-rate-limiting) — Rate limiting throttles and retries provider calls so a burst or an HTTP 429 does not fail the rows that could have waited.
- [What is structured output?](https://pixeltable.com/learn/what-is-structured-output) — Structured output constrains a model to JSON or a schema, so downstream columns get typed fields instead of free prose.
- [What is a model provider?](https://pixeltable.com/learn/what-is-a-model-provider) — A model provider is the API that serves an embedding, a transcription, a vision model, or a chat model a column can call.
- [What is local inference?](https://pixeltable.com/learn/what-is-local-inference) — Local inference runs a model on hardware you control, instead of calling a hosted model API.
- [What is an LLM framework?](https://pixeltable.com/learn/what-is-an-llm-framework) — An LLM framework is a library for chaining model calls, tools, and prompts in application code.

## Storage

- [What is a media store?](https://pixeltable.com/learn/what-is-a-media-store) — A media store holds the bytes for images, video, audio, and documents. The catalog stores a typed pointer, not the bytes themselves.
- [What is object storage?](https://pixeltable.com/learn/what-is-object-storage) — Object storage is a bucket API that holds blobs addressed by key. A database may point at those keys instead of copying every byte.
- [What is a local catalog?](https://pixeltable.com/learn/what-is-a-local-catalog) — A local catalog is the on-disk database, caches, and media home where tables live when you run the engine on your own machine.
- [What is bring-your-own-bucket?](https://pixeltable.com/learn/what-is-bring-your-own-bucket) — Bring-your-own-bucket means media URIs point at a bucket you already pay for, so the catalog stores typed pointers instead of a second copy of the bytes.
- [What is a file cache?](https://pixeltable.com/learn/what-is-a-file-cache) — A file cache keeps a local copy of media bytes so repeated transforms do not download the same object again.
- [What is lakehouse export?](https://pixeltable.com/learn/what-is-lakehouse-export) — Lakehouse export ships table data, often as Iceberg or Arrow, into an analytics lake without making the lake the system of record for media pipelines.
- [What is a data warehouse?](https://pixeltable.com/learn/what-is-a-warehouse) — A data warehouse stores analytical tables for SQL over large scans. Unstructured files are usually paths or external tables, not typed media with model columns.

## Versioning

- [What is time travel?](https://pixeltable.com/learn/what-is-time-travel) — Time travel queries a prior version of a table, so you can see rows and computed outputs as they were after an earlier change.
- [What is data lineage?](https://pixeltable.com/learn/what-is-data-lineage) — Data lineage traces an output back to the source rows, the model, and the column definition that produced it.
- [What is dataset versioning?](https://pixeltable.com/learn/what-is-dataset-versioning) — Dataset versioning pins which rows, labels, and derived features a training or eval run used.
- [What is reproducibility in a data pipeline?](https://pixeltable.com/learn/what-is-reproducibility) — Reproducibility means you can inspect the same inputs and column definitions and get the same cached outputs, or a known diff.

## Serving

- [What is HTTP serving for tables?](https://pixeltable.com/learn/what-is-http-serving) — HTTP serving exposes insert, query, and computed results as routes from the same schema that defines the tables.
- [What is a query decorator?](https://pixeltable.com/learn/what-is-a-query-decorator) — A query decorator names a reusable query you can call from an agent, from HTTP, or from another column.
- [How does the local-to-cloud loop work?](https://pixeltable.com/learn/what-is-the-local-cloud-loop) — The local-to-cloud loop runs the same schema on a laptop catalog and on a hosted database, so you change it locally and then apply it remotely.
- [What is a CLI catalog?](https://pixeltable.com/learn/what-is-a-cli-catalog) — A CLI catalog is the same tables and services, operated from the terminal, that a dashboard can also show.
- [What is a multimodal backend?](https://pixeltable.com/learn/what-is-a-multimodal-backend) — A multimodal backend is storage, orchestration, retrieval, and HTTP for media and models in one deployable schema.

## Evaluation

- [What is ML evaluation?](https://pixeltable.com/learn/what-is-ml-evaluation) — ML evaluation scores model or pipeline outputs against labels or a judge, so you can compare versions.
- [What is a feature store?](https://pixeltable.com/learn/what-is-a-feature-store) — A feature store serves consistent numeric or categorical features to training and to inference, usually for tabular models.
- [What is active learning?](https://pixeltable.com/learn/what-is-active-learning) — Active learning chooses which unlabeled examples a person should label next, so the model improves faster than if you labeled at random.
- [What is semantic deduplication?](https://pixeltable.com/learn/what-is-semantic-deduplication) — Semantic deduplication drops near-duplicate items by embedding similarity, not only by exact hash equality.
- [What is dataset curation?](https://pixeltable.com/learn/what-is-dataset-curation) — Dataset curation selects, cleans, and versions the rows that train or evaluate a model.
