---
title: "What is document extraction?"
description: "Document extraction pulls text, layout, or tables out of a file so later columns can chunk, embed, or summarize."
url: "https://pixeltable.com/learn/what-is-document-extraction"
updated: "2026-09-29"
vertical: "Documents"
doc: "https://docs.pixeltable.com/datastore/computed-columns"
---

# What is document extraction?

Document extraction pulls text, layout, or tables out of a file so later columns can chunk, embed, or summarize.

Updated: 2026-09-29
Part of [What is a document column?](https://pixeltable.com/learn/what-is-a-document-column).

## On this page

- [How it works](https://pixeltable.com/learn/what-is-document-extraction#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-document-extraction#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-document-extraction#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-document-extraction#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-document-extraction#questions)

## How it works {#how-it-works}


- Open the file as a document.
- Write text or structure into columns.
- Downstream columns chunk or answer from that text.

## What it is not {#what-it-is-not}

It is not a summary with no extractable text left behind, and it is not treating the PDF as an image-only blob.

## document extraction: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Keeps | Text you can index | Only a summary |
| Layout | Optional, when the extractor provides it | Required for every file |
| Source | Still the document row | Discarded after extraction |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable extraction is a computed column on pxt.Document. Chunking and Q&A are further columns or views.

## Questions {#questions}

### How does document extraction work? {#faq-1}

Open the file as a document. Write text or structure into columns. Downstream columns chunk or answer from that text.

### What is document extraction often confused with? {#faq-2}

It is not a summary with no extractable text left behind, and it is not treating the PDF as an image-only blob.

## In the blog

- [Claude 3.5 Sonnet Vision: Building Multimodal AI Applications with Anthropic's Claude and Pixeltable](https://pixeltable.com/blog/claude-anthropic-multimodal-integration-pixeltable)

## Related

- [Documentation](https://docs.pixeltable.com/datastore/computed-columns)
- [Build Production Document RAG Pipelines with Pixeltable](https://pixeltable.com/blog/document-pdf-processing-rag-pixeltable)
- [Production RAG](https://pixeltable.com/use-cases/production-rag-implementation)
- [pdf-to-text](https://pixeltable.com/tools/pdf-to-text)
- [document-summarizer](https://pixeltable.com/tools/document-summarizer)
