Retrieval

What is chunking?

Chunking splits a source into retrieval units — sentences, pages, token windows — small enough to embed and cite.

Updated

How it works

  • Choose a unit that fits the question you expect.
  • Each unit becomes a row that points at the source.
  • The embedding is computed on the unit, not on the entire file.

What it is not

It is not embedding the whole file as one vector, and it is not throwing away the source after the split.

chunking: this, and the thing it is confused with

chunking: this, and the thing it is confused with
ThisNot this
UnitA passage you can quoteThe whole PDF as one vector
ParentKeptDiscarded after the split
Too-large unitsMiss the local factA window that still cites a page

Where Pixeltable fits

Pixeltable chunks with an iterator view, so each passage row still points at the document or string it came from.

Questions

How does chunking work?
Choose a unit that fits the question you expect. Each unit becomes a row that points at the source. The embedding is computed on the unit, not on the entire file.
What is chunking often confused with?
It is not embedding the whole file as one vector, and it is not throwing away the source after the split.