---
title: "What is object detection?"
description: "Object detection finds instances of objects in an image or frame and returns boxes, usually with a class and a score."
url: "https://pixeltable.com/learn/what-is-object-detection"
updated: "2026-09-29"
vertical: "Vision"
doc: "https://docs.pixeltable.com/howto/cookbooks/video/video-extract-frames"
---

# What is object detection?

Object detection finds instances of objects in an image or frame and returns boxes, usually with a class and a score.

Updated: 2026-09-29
Part of [What is computer vision?](https://pixeltable.com/learn/what-is-computer-vision).

## On this page

- [How it works](https://pixeltable.com/learn/what-is-object-detection#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-object-detection#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-object-detection#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-object-detection#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-object-detection#questions)

## How it works {#how-it-works}


- The model reads a frame.
- It writes boxes and labels.
- Those labels are stored on the frame row.

## What it is not {#what-it-is-not}

It is not classification of the whole picture, and it is not semantic search without boxes.

## object detection: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Output | Boxes and classes | One label for the whole image |
| Location | Where in the frame | Only that something exists |
| Search | Frames that contain a class | A similar picture with no box |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable assigns a detector such as YOLOX on the frame column of a video view.

## Questions {#questions}

### How does object detection work? {#faq-1}

The model reads a frame. It writes boxes and labels. Those labels are stored on the frame row.

### What is object detection often confused with? {#faq-2}

It is not classification of the whole picture, and it is not semantic search without boxes.

## In the blog

- [Rerun vs Pixeltable: From 450 Lines to 15 in Computer Vision Pipelines](https://pixeltable.com/blog/rerun-vs-pixeltable-computer-vision)
- [Pixeltable Supports the CV Community with a Maintained YOLOX Fork](https://pixeltable.com/blog/pixeltable-yolox-fork)
- [YOLOX Object Detection for Video Analysis: Complete Guide with Pixeltable](https://pixeltable.com/blog/object-detection-videos-yolox)

## Related

- [Documentation](https://docs.pixeltable.com/howto/cookbooks/video/video-extract-frames)
- [YOLOX Object Detection for Video Analysis: Complete Guide with Pixeltable](https://pixeltable.com/blog/object-detection-videos-yolox)
- [Computer vision pipeline](https://pixeltable.com/use-cases/computer-vision-pipeline-optimization)
- [Pixeltable vs Voxel51](https://pixeltable.com/compare/pixeltable-vs-voxel51)
