---
title: "What is cross-modal search?"
description: "Cross-modal search queries one modality with another — a sentence against frames, a still against a video library — because they share an embedding space."
url: "https://pixeltable.com/learn/what-is-cross-modal-search"
updated: "2026-09-29"
vertical: "Retrieval"
doc: "https://docs.pixeltable.com/datastore/embedding-index"
---

# What is cross-modal search?

Cross-modal search queries one modality with another — a sentence against frames, a still against a video library — because they share an embedding space.

Updated: 2026-09-29
Part of [What is semantic search?](https://pixeltable.com/learn/what-is-semantic-search).

## On this page

- [How it works](https://pixeltable.com/learn/what-is-cross-modal-search#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-cross-modal-search#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-cross-modal-search#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-cross-modal-search#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-cross-modal-search#questions)

## How it works {#how-it-works}


- One model places text and images in the same space.
- The query is embedded as text or as an image.
- The nearest frames or pictures are the hits, with their source rows.

## What it is not {#what-it-is-not}

It is not transcript-only search, and it is not filename search.

## cross-modal search: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Query | Text or a still | Only the same file type, by name |
| Index | Frames or images in that shared space | A transcript |
| Hit | A picture that matches the phrase | A document that contains the words |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable uses a shared embedding such as CLIP on a frame or image column, then similarity with a string or an image.

## Questions {#questions}

### How does cross-modal search work? {#faq-1}

One model places text and images in the same space. The query is embedded as text or as an image. The nearest frames or pictures are the hits, with their source rows.

### What is cross-modal search often confused with? {#faq-2}

It is not transcript-only search, and it is not filename search.

## In the blog

- [Build Your Own Multimodal Search Engine with Pixeltable](https://pixeltable.com/blog/pixelsearch-multimodal-search-engine)
- [Find a Video from an Image with Pixeltable](https://pixeltable.com/blog/find-video-from-image-pixelsearch)

## Related

- [Documentation](https://docs.pixeltable.com/datastore/embedding-index)
- [Find a Video from an Image with Pixeltable](https://pixeltable.com/blog/find-video-from-image-pixelsearch)
- [ClipFinder: Natural-Language Video Moment Search in One app.py](https://pixeltable.com/blog/clipfinder-semantic-video-moment-search)
- [Multimodal application](https://pixeltable.com/use-cases/multimodal-ai-application-development)
