Documents
What is document extraction?
Document extraction pulls text, layout, or tables out of a file so later columns can chunk, embed, or summarize.
Updated · Part of What is a document column?
How it works
- Open the file as a document.
- Write text or structure into columns.
- Downstream columns chunk or answer from that text.
What it is not
It is not a summary with no extractable text left behind, and it is not treating the PDF as an image-only blob.
document extraction: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Keeps | Text you can index | Only a summary |
| Layout | Optional, when the extractor provides it | Required for every file |
| Source | Still the document row | Discarded after extraction |
Where Pixeltable fits
Pixeltable extraction is a computed column on pxt.Document. Chunking and Q&A are further columns or views.
Questions
- How does document extraction work?
- Open the file as a document. Write text or structure into columns. Downstream columns chunk or answer from that text.
- What is document extraction often confused with?
- It is not a summary with no extractable text left behind, and it is not treating the PDF as an image-only blob.