---
title: "What is a document splitter?"
description: "A document splitter cuts a document into retrieval units — sentences, paragraphs, or pages — as rows over the source file."
url: "https://pixeltable.com/learn/what-is-a-document-splitter"
updated: "2026-09-29"
vertical: "Documents"
doc: "https://docs.pixeltable.com/datastore/computed-columns"
---

# What is a document splitter?

A document splitter cuts a document into retrieval units — sentences, paragraphs, or pages — as rows over the source file.

Updated: 2026-09-29
Part of [What is chunking?](https://pixeltable.com/learn/what-is-chunking).

## On this page

- [How it works](https://pixeltable.com/learn/what-is-a-document-splitter#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-a-document-splitter#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-a-document-splitter#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-a-document-splitter#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-a-document-splitter#questions)

## How it works {#how-it-works}


- The view iterates the document column.
- Each unit is a row with the text and a pointer to the file.
- A changed file rebuilds the units that came from it.

## What it is not {#what-it-is-not}

It is not a one-off split in a notebook that never updates when the PDF changes.

## document splitter: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Input | A document column | A string you split once |
| Output | Passage rows | A list in memory |
| On edit | Units follow the file | You rerun the notebook |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable’s view sets iterator=document_splitter(...) on a pxt.Document column.

```python
import pixeltable as pxt
from pixeltable.functions.document import document_splitter

TableModel = pxt.model_base()

class Docs(TableModel, name='docs'):
    document: pxt.Document

class Chunks(
    TableModel,
    name='chunks',
    base=Docs,
    iterator=document_splitter(document=Docs.document, separators='sentence'),
):
    pass
```

## Questions {#questions}

### How does document splitter work? {#faq-1}

The view iterates the document column. Each unit is a row with the text and a pointer to the file. A changed file rebuilds the units that came from it.

### What is document splitter often confused with? {#faq-2}

It is not a one-off split in a notebook that never updates when the PDF changes.

## In the blog

- [Build Production Document RAG Pipelines with Pixeltable](https://pixeltable.com/blog/document-pdf-processing-rag-pixeltable)

## Related

- [Documentation](https://docs.pixeltable.com/datastore/computed-columns)
- [Build Production Document RAG Pipelines with Pixeltable](https://pixeltable.com/blog/document-pdf-processing-rag-pixeltable)
- [Iterate on Your Data, Not Your Infrastructure: The Multimodal Experimentation Loop](https://pixeltable.com/blog/iterate-on-data-not-infrastructure)
- [Production RAG](https://pixeltable.com/use-cases/production-rag-implementation)
