Documents

What is OCR?

OCR, optical character recognition, turns pixels of text — scans, screenshots, photos — into strings.

Updated

How it works

  • The input is an image or a scanned page.
  • A model reads the glyphs and writes text.
  • That text can then be chunked and searched like any other string.

What it is not

It is not extracting text already embedded in a digital PDF, and it is not captioning a scene.

OCR: this, and the thing it is confused with

OCR: this, and the thing it is confused with
ThisNot this
InputPixels of writingA digital PDF with a text layer
OutputCharactersA description of the picture
Next stepIndex the stringStop at the image

Where Pixeltable fits

Pixeltable runs OCR as a computed column on an image or document. The string it writes is what later search uses.

Questions

How does OCR work?
The input is an image or a scanned page. A model reads the glyphs and writes text. That text can then be chunked and searched like any other string.
What is OCR often confused with?
It is not extracting text already embedded in a digital PDF, and it is not captioning a scene.