---
title: "What is chunking?"
description: "Chunking splits a source into retrieval units — sentences, pages, token windows — small enough to embed and cite."
url: "https://pixeltable.com/learn/what-is-chunking"
updated: "2026-09-29"
vertical: "Retrieval"
doc: "https://docs.pixeltable.com/datastore/computed-columns"
---

# What is chunking?

Chunking splits a source into retrieval units — sentences, pages, token windows — small enough to embed and cite.

Updated: 2026-09-29


## On this page

- [How it works](https://pixeltable.com/learn/what-is-chunking#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-chunking#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-chunking#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-chunking#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-chunking#questions)

## How it works {#how-it-works}


- Choose a unit that fits the question you expect.
- Each unit becomes a row that points at the source.
- The embedding is computed on the unit, not on the entire file.

## What it is not {#what-it-is-not}

It is not embedding the whole file as one vector, and it is not throwing away the source after the split.

## chunking: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Unit | A passage you can quote | The whole PDF as one vector |
| Parent | Kept | Discarded after the split |
| Too-large units | Miss the local fact | A window that still cites a page |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable chunks with an iterator view, so each passage row still points at the document or string it came from.

## Questions {#questions}

### How does chunking work? {#faq-1}

Choose a unit that fits the question you expect. Each unit becomes a row that points at the source. The embedding is computed on the unit, not on the entire file.

### What is chunking often confused with? {#faq-2}

It is not embedding the whole file as one vector, and it is not throwing away the source after the split.

## Related

- [Documentation](https://docs.pixeltable.com/datastore/computed-columns)
- [Build Production Document RAG Pipelines with Pixeltable](https://pixeltable.com/blog/document-pdf-processing-rag-pixeltable)
- [Production RAG Systems: Building Data-Centric RAG Applications at Scale](https://pixeltable.com/blog/production-rag-data-centric)
- [Production RAG](https://pixeltable.com/use-cases/production-rag-implementation)
