---
title: "What is CLIP?"
description: "CLIP is a model that maps images and text into one space, so a sentence can retrieve a matching picture or frame."
url: "https://pixeltable.com/learn/what-is-clip"
updated: "2026-09-29"
vertical: "Vision"
doc: "https://docs.pixeltable.com/datastore/embedding-index"
---

# What is CLIP?

CLIP is a model that maps images and text into one space, so a sentence can retrieve a matching picture or frame.

Updated: 2026-09-29


## On this page

- [How it works](https://pixeltable.com/learn/what-is-clip#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-clip#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-clip#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-clip#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-clip#questions)

## How it works {#how-it-works}


- The same model embeds a phrase and a picture.
- Nearby vectors are the matches.
- It does not, by itself, draw boxes around objects.

## What it is not {#what-it-is-not}

It is not a captioning language model, and it is not a text-only sentence transformer.

## CLIP: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Query | Text or an image | Only keywords in a filename |
| Space | Shared by text and images | Text-only, or image-only |
| Does not do | Boxes and classes | Nearest picture |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable plugs CLIP in as the embedding function on an image or frame column.

## Questions {#questions}

### How does CLIP work? {#faq-1}

The same model embeds a phrase and a picture. Nearby vectors are the matches. It does not, by itself, draw boxes around objects.

### What is CLIP often confused with? {#faq-2}

It is not a captioning language model, and it is not a text-only sentence transformer.

## In the blog

- [LectureSync: Slide-Grounded Lecture Q&A Without an Aligner Service](https://pixeltable.com/blog/lecturesync-slide-lecture-qa)
- [AdRadar: Creative Fatigue and Copycat Search Without a Tagging Farm](https://pixeltable.com/blog/adradar-creative-fatigue-visual-search)
- [SnapCatalog: Shop-the-Look Search Without a Tagging Farm](https://pixeltable.com/blog/snapcatalog-product-visual-search)
- [ClipFinder: Natural-Language Video Moment Search in One app.py](https://pixeltable.com/blog/clipfinder-semantic-video-moment-search)
- [Find a Video from an Image with Pixeltable](https://pixeltable.com/blog/find-video-from-image-pixelsearch)
- [VideoRAG: Where Pixeltable Stores Frames and Indexes](https://pixeltable.com/blog/videorag-where-pixeltable-stores-indexes)

## Related

- [Documentation](https://docs.pixeltable.com/datastore/embedding-index)
- [Find a Video from an Image with Pixeltable](https://pixeltable.com/blog/find-video-from-image-pixelsearch)
- [ClipFinder: Natural-Language Video Moment Search in One app.py](https://pixeltable.com/blog/clipfinder-semantic-video-moment-search)
- [Video intelligence pipeline](https://pixeltable.com/use-cases/video-content-analysis)
- [Computer vision pipeline](https://pixeltable.com/use-cases/computer-vision-pipeline-optimization)
- [semantic-search](https://pixeltable.com/tools/semantic-search)
