Documents
What is OCR?
OCR, optical character recognition, turns pixels of text — scans, screenshots, photos — into strings.
Updated
How it works
- The input is an image or a scanned page.
- A model reads the glyphs and writes text.
- That text can then be chunked and searched like any other string.
What it is not
It is not extracting text already embedded in a digital PDF, and it is not captioning a scene.
OCR: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Input | Pixels of writing | A digital PDF with a text layer |
| Output | Characters | A description of the picture |
| Next step | Index the string | Stop at the image |
Where Pixeltable fits
Pixeltable runs OCR as a computed column on an image or document. The string it writes is what later search uses.
Questions
- How does OCR work?
- The input is an image or a scanned page. A model reads the glyphs and writes text. That text can then be chunked and searched like any other string.
- What is OCR often confused with?
- It is not extracting text already embedded in a digital PDF, and it is not captioning a scene.