Vision

What is CLIP?

CLIP is a model that maps images and text into one space, so a sentence can retrieve a matching picture or frame.

Updated

How it works

  • The same model embeds a phrase and a picture.
  • Nearby vectors are the matches.
  • It does not, by itself, draw boxes around objects.

What it is not

It is not a captioning language model, and it is not a text-only sentence transformer.

CLIP: this, and the thing it is confused with

CLIP: this, and the thing it is confused with
ThisNot this
QueryText or an imageOnly keywords in a filename
SpaceShared by text and imagesText-only, or image-only
Does not doBoxes and classesNearest picture

Where Pixeltable fits

Pixeltable plugs CLIP in as the embedding function on an image or frame column.

Questions

How does CLIP work?
The same model embeds a phrase and a picture. Nearby vectors are the matches. It does not, by itself, draw boxes around objects.
What is CLIP often confused with?
It is not a captioning language model, and it is not a text-only sentence transformer.