Video

How does video search work?

Video search finds a moment inside a clip, not only a filename. The usual path is to sample frames, embed each frame, and rank those frames by similarity to a text phrase or a still image.

Updated

What it is

A video file is a sequence of pictures plus a timeline. Search that returns “beach.mp4” has matched a title. Search that returns a frame at 00:01:12 has matched the picture. The second kind is what people mean by video search, video retrieval, or finding a moment.

How it works

Three steps. The frame rate you choose is the tradeoff: more frames catch shorter moments and cost more embedding work.

  • Sample. A view walks the video at a fixed rate, one row per frame, and keeps the frame index and the timestamp.
  • Embed. A model such as CLIP turns each frame into a vector in a space where a text phrase and a picture can be neighbors.
  • Query. The question is embedded the same way. The nearest frames come back with the source video and the time offset.

What it is not

Sending the whole file to a chat model on every question is not an index. It can describe one clip. It does not rank a library, and it does not get cheaper as the library grows.

A transcript search finds words that were spoken. It misses a red car that nobody named. Visual search and speech search answer different questions. A library often wants both, as two indexes on the same clip.

What you keep at each step of video search

What you keep at each step of video search
StepWhat happensWhat you can return
Sample framesOne row per frame at a chosen rateFrame index and timestamp on the source video
Embed framesA model maps each frame to a vectorNeighbors of a text phrase or a still image
QueryRank frames by similarityThe moment, not only the filename

Where Pixeltable fits

In Pixeltable a frame view uses frame_iterator on a Video column. An embedding index on the frame column is what the similarity query reads. The parent video and the timestamp stay on the frame row, so the hit is a moment. The step-by-step version is the video intelligence use case.

import pixeltable as pxt
from pixeltable.functions.huggingface import clip
from pixeltable.functions.video import frame_iterator
TableModel = pxt.model_base()
image_embed = clip.using(model_id='openai/clip-vit-base-patch32')
class Videos(TableModel, name='videos'):
video: pxt.Video
class Frames(
TableModel,
name='frames',
base=Videos,
iterator=frame_iterator(video=Videos.video, fps=1),
):
__indexes__ = [
pxt.EmbeddingIndex(frame, image_embed=image_embed),
]
@pxt.query
def find_moment(query: str, limit: int = 5):
sim = Frames.frame.similarity(string=query)
return (
Frames.order_by(sim, asc=False)
.limit(limit)
.select(Frames.video, Frames.frame_idx, Frames.pos_msec, sim)
)

Questions

How many frames should I index?
One frame per second is a common start. Higher rates catch shorter actions and create more rows to embed. Lower rates are cheaper and miss brief moments. The right rate is the shortest event you still need to find.
Can I search video with a sentence?
Yes, if the embedding model places text and images in one space. CLIP-style models do. The sentence and the frame are both vectors, and similarity ranks frames against the sentence.
Is a transcript enough?
A transcript finds what was said. It does not find a visual match that nobody described. Use a transcript index for speech and a frame index for pictures.