Documents

What is document extraction?

Document extraction pulls text, layout, or tables out of a file so later columns can chunk, embed, or summarize.

Updated · Part of What is a document column?

How it works

  • Open the file as a document.
  • Write text or structure into columns.
  • Downstream columns chunk or answer from that text.

What it is not

It is not a summary with no extractable text left behind, and it is not treating the PDF as an image-only blob.

document extraction: this, and the thing it is confused with

document extraction: this, and the thing it is confused with
ThisNot this
KeepsText you can indexOnly a summary
LayoutOptional, when the extractor provides itRequired for every file
SourceStill the document rowDiscarded after extraction

Where Pixeltable fits

Pixeltable extraction is a computed column on pxt.Document. Chunking and Q&A are further columns or views.

Questions

How does document extraction work?
Open the file as a document. Write text or structure into columns. Downstream columns chunk or answer from that text.
What is document extraction often confused with?
It is not a summary with no extractable text left behind, and it is not treating the PDF as an image-only blob.