---
title: "Computer Vision Pipeline: Object Detection, Classification, and Search"
description: "Build optimized computer vision workflows with Pixeltable. Run YOLOX, CLIP, and custom models as computed columns with automatic batching, caching, and incremental processing."
keywords:
  - computer vision pipeline
  - object detection python
  - YOLOX pipeline
  - CLIP image search
  - cv data pipeline
  - image classification automation
  - computer vision at scale
complexity: "intermediate"
estimated_time: "30 min"
url: "https://pixeltable.com/use-cases/computer-vision-pipeline-optimization"
---

# Computer Vision Pipeline: Object Detection, Classification, and Search

Build optimized computer vision workflows with Pixeltable. Run YOLOX, CLIP, and custom models as computed columns with automatic batching, caching, and incremental processing.

## Prerequisites

- Familiarity with computer vision concepts
- Python and basic ML experience

## The Problem

Computer vision pipelines require managing image preprocessing, running multiple models (detection, classification, embedding), storing results, and keeping everything in sync. Adding a new model or changing thresholds means reprocessing millions of images manually.

## The Solution

Pixeltable makes CV pipelines declarative. Define models as computed columns, and Pixeltable handles execution, batching, caching, and incremental updates. Change a threshold and only affected results are recomputed.

## Implementation

### Image Table

Create a table for images with metadata tracking.

```python
import pixeltable as pxt

# Create image processing pipeline
images = pxt.create_table('app.images', {
    'image': pxt.Image,
    'source': pxt.String,
    'timestamp': pxt.Timestamp,
    'camera_id': pxt.String,
})

# Insert images: local, S3, or URL
images.insert([
    {'image': 's3://bucket/cam01/frame_001.jpg',
     'source': 'warehouse', 'camera_id': 'cam-01'},
    {'image': '/data/inspection/part_42.png',
     'source': 'qc-station', 'camera_id': 'cam-02'},
])
```

Images are stored as native types with full metadata. No separate image registry or file management.


### Object Detection

Run YOLOX object detection on every image automatically.

```python
from pixeltable.functions import yolox

# Object detection as a computed column
images.add_computed_column(
    detections=yolox(
        images.image,
        model_id='yolox_m',
        threshold=0.6
    )
)

# Extract detection counts for filtering
images.add_computed_column(
    num_objects=images.detections.apply(lambda d: len(d['labels']))
)

# Filter to images with specific objects
busy_frames = images.select(
    images.image, images.detections
).where(
    images.num_objects > 5
).collect()
```

Detection runs automatically for every image, existing and future. Change the threshold and only affected rows recompute.


### Visual Embeddings

Generate CLIP embeddings for visual similarity search.

```python
from pixeltable.functions.huggingface import clip

# CLIP embeddings for visual search
images.add_embedding_index(
    'image',
    image_embed=clip.using(
        model_id='openai/clip-vit-base-patch32'
    )
)

# Find visually similar images
similar = images.select(
    images.image, images.detections, images.source
).order_by(
    images.image.similarity(string='forklift in warehouse'),
    asc=False
).limit(10)

# Or search by reference image
ref_image = '/data/reference/defect_sample.png'
matches = images.select(
    images.image, images.source
).order_by(
    images.image.similarity(ref_image), asc=False
).limit(20)
```

CLIP enables text-to-image and image-to-image search in the same index. No separate vector DB needed.


### Custom Models

Add your own models as user-defined functions.

```python
import torch
from PIL import Image as PILImage

# Custom classification model as a UDF
@pxt.udf
def classify_defect(image: PILImage.Image) -> dict:
    """Run custom defect classifier on an image."""
    model = torch.load('models/defect_classifier.pt')
    model.eval()
    tensor = preprocess(image)
    with torch.no_grad():
        pred = model(tensor.unsqueeze(0))
    return {
        'class': LABELS[pred.argmax().item()],
        'confidence': pred.max().item()
    }

# Add as computed column: runs on every image
images.add_computed_column(
    defect_result=classify_defect(images.image)
)

# Filter by classification results
defects = images.select(
    images.image, images.defect_result
).where(
    images.defect_result['confidence'] > 0.9
).collect()
```

Any Python function becomes a computed column. Pixeltable handles execution, caching, and error recovery.


## Benefits

- 10x faster CV pipeline development vs custom code
- Automatic batching and GPU utilization
- Built-in result caching eliminates redundant inference
- Incremental processing: only new images trigger computation
- Mix vendor models (YOLOX, CLIP) with custom PyTorch models seamlessly

## Use Cases

- Manufacturing quality control and defect detection
- Retail inventory monitoring and shelf analytics
- Security and surveillance image analysis
- Medical image classification and triage
- Autonomous vehicle perception data processing

## Performance


| Metric | Value | Description |

| --- | --- | --- |

| Development Speed | 10x faster | vs building custom CV infrastructure |

| Compute Savings | 70% | With caching and incremental processing |

## Requirements

- Python 3.9+
- GPU recommended for large-scale inference
- PyTorch for custom model support

## Resources

- [YOLOX Object Detection in Videos](https://pixeltable.com/blog/object-detection-videos-yolox) - Detailed YOLOX integration guide
- [CV Data Curation & Annotation](https://pixeltable.com/blog/automate-cv-data-pixeltable) - FiftyOne and Label Studio integration
- [Accelerating Multimodal Annotations](https://pixeltable.com/blog/accelerating-multimodal-ai-data-annotations-pixeltable) - Speed up annotation workflows with AI