---
title: "Rerun vs Pixeltable: From 450 Lines to 15 in Computer Vision Pipelines"
date: "2025-12-16"
author: "Pixeltable Team"
tags:
  - Computer Vision
  - Object Detection
  - Video Processing
  - Rerun Alternative
  - Declarative AI
  - DETR
  - Video Analysis
description: "Compare Rerun's real-time streaming visualization with Pixeltable's declarative batch processing for video object detection. See why declarative approaches eliminate object tracking code entirely."
url: "https://pixeltable.com/blog/rerun-vs-pixeltable-computer-vision"
---

# Rerun vs Pixeltable: From 450 Lines to 15 in Computer Vision Pipelines

Rerun's [detect_and_track_objects](https://github.com/rerun-io/rerun/tree/main/examples/python/detect_and_track_objects) example is an impressive 450-line Python script that detects and tracks objects in video. It uses DETR for detection, CSRT for tracking, and Rerun for visualization.

 
We can achieve the same result in Pixeltable with about 15 lines of code. Here's why, and what it reveals about the difference between streaming and declarative approaches to computer vision.

 
## The Rerun Approach (450 Lines)

 
```python
# Simplified pseudocode of Rerun's approach
cap = cv2.VideoCapture(video_path)
detector = Detector() # DETR model
trackers = []

while cap.isOpened():
 ret, frame = cap.read()
 rr.set_time("frame", sequence=frame_idx)
 
 # Run expensive detection every 40 frames
 if frame_idx % 40 == 0:
 detections = detector.detect(frame)
 trackers = update_trackers_with_detections(trackers, detections, frame)
 else:
 # Run cheap tracking every frame
 for tracker in trackers:
 tracker.update(frame)
 
 # Log to Rerun viewer
 for tracker in trackers:
 rr.log(f"video/tracked/{tracker.id}", rr.Boxes2D(...))
 
 frame_idx += 1
```

 
**Why does Rerun need tracking?** DETR detection is expensive (~100ms per frame). Running it on every frame of a 30fps video would take forever. So Rerun runs detection every 40 frames and uses cheap CSRT tracking (~1ms) to interpolate object positions in between.

 
## The Pixeltable Approach (15 Lines)

 
```python
import pixeltable as pxt
from pixeltable.functions.video import frame_iterator
from pixeltable.functions.huggingface import detr_for_object_detection
import pixeltable.functions as pxtf

# Create table and view
pxt.create_dir('demo', if_exists='replace_force')
videos = pxt.create_table('demo.videos', {'video': pxt.Video})
frames = pxt.create_view('demo.frames', videos, iterator=frame_iterator(videos.video, fps=5))

# Insert video and add detection
videos.insert([{'video': 'path/to/video.mp4'}])
frames.add_computed_column(detections=detr_for_object_detection(frames.frame, model_id='facebook/detr-resnet-50'))

# Output annotated video
frames.group_by(videos).select(
 pxt.functions.video.make_video(frames.pos, pxtf.vision.draw_bounding_boxes(frames.frame, frames.detections.boxes))
).show()
```

 
**Why doesn't Pixeltable need tracking?** Because we're not streaming. We extract frames at 5 fps (not 30), run detection on every extracted frame, and reassemble them. No interpolation needed.

 
## What's Superfluous in Pixeltable?

 
### 1. Object Tracking (CSRT)

 

 - **Rerun needs it:** To maintain object IDs between sparse detections

 - **Pixeltable doesn't:** Every frame has its own complete detection results

 

 
### 2. Manual Frame Loop

 

 - **Rerun needs it:** `while cap.isOpened(): ret, frame = cap.read()`

 - **Pixeltable doesn't:** `frame_iterator` handles extraction declaratively

 

 
### 3. Stateful Tracker Management

 

 - **Rerun needs it:** `trackers = update_trackers_with_detections(...)`, handling tracker creation/deletion

 - **Pixeltable doesn't:** Each frame is independent, with no state to manage

 

 
### 4. Manual Visualization Code

 

 - **Rerun needs it:** `rr.log(f"video/tracked/{tracker.id}", rr.Boxes2D(...))`

 - **Pixeltable doesn't:** `draw_bounding_boxes()` + `make_video()` handles it

 

 
## The Architectural Difference

 
| Aspect | Rerun (Streaming) | Pixeltable (Declarative) |
| --- | --- | --- |
| Processing | Frame-by-frame loop | Computed columns on extracted frames |
| Detection Frequency | Every N frames (expensive) | Every extracted frame (at lower fps) |
| Tracking | Required (maintain IDs between detections) | Not needed (each frame independent) |
| State Management | Manual (tracker lifecycle) | Automatic (table handles it) |
| Results | Ephemeral (streaming) | Persistent (queryable) |
| Incremental Updates | Rerun entire script | Automatic (add new videos) |

 
## When to Use Each

 
### Use Rerun When:

 

 - You need real-time visualization during development

 - You want interactive timeline scrubbing

 - You're debugging a computer vision pipeline

 - You need 3D visualization

 

 
### Use Pixeltable When:

 

 - You're batch processing video datasets

 - You want to query detection results later

 - You need to compare multiple models

 - You want [incremental updates](/blog/incremental-embedding-indexes) as new videos arrive

 - You're building a production pipeline

 

 
## The Code Comparison

 
| Component | Rerun | Pixeltable |
| --- | --- | --- |
| Frame extraction | ~20 lines | 1 line (frame_iterator) |
| Detection | ~50 lines (Detector class) | 1 line (detr_for_object_detection) |
| Tracking | ~150 lines (Tracker class + update logic) | 0 lines (not needed) |
| Visualization | ~30 lines | 1 line (draw_bounding_boxes) |
| Video output | N/A (streaming viewer) | 1 line (make_video) |
| Total | ~450 lines | ~15 lines |

 
## Try It Yourself

 
```python
%pip install pixeltable torch transformers

import pixeltable as pxt
from pixeltable.functions.video import frame_iterator
from pixeltable.functions.huggingface import detr_for_object_detection
import pixeltable.functions as pxtf

pxt.create_dir('demo', if_exists='replace_force')
videos = pxt.create_table('demo.videos', {'video': pxt.Video})
frames = pxt.create_view('demo.frames', videos, iterator=frame_iterator(videos.video, fps=5))

videos.insert([{'video': 'https://raw.githubusercontent.com/pixeltable/pixeltable/release/docs/resources/bangkok.mp4'}])
frames.add_computed_column(detections=detr_for_object_detection(frames.frame, model_id='facebook/detr-resnet-50'))

# View annotated video
frames.group_by(videos).select(
 pxt.functions.video.make_video(frames.pos, pxtf.vision.draw_bounding_boxes(frames.frame, frames.detections.boxes))
).show()
```

 
## Bonus: Panoptic Segmentation

 
The Rerun example also uses panoptic segmentation. While not needed for basic object detection, segmentation provides richer scene understanding. Here's how to use it in Pixeltable.

 
### What is Panoptic Segmentation?

 
Panoptic segmentation unifies two tasks:

 

 - **Semantic segmentation:** Classify every pixel (sky, road, grass)

 - **Instance segmentation:** Identify individual objects (car #1, car #2, person #1)

 

 
The result is a complete scene understanding where every pixel belongs to either a "stuff" class (amorphous regions like sky) or a "thing" instance (countable objects like cars).

 
### The Segmentation Data

 
DETR's panoptic model returns:

 
```python
{
 'segmentation': tensor, # H×W tensor where each pixel = segment ID
 'segments_info': [
 {'id': 1, 'label_id': 0, 'score': 0.98, 'area': 45000}, # sky
 {'id': 2, 'label_id': 1, 'score': 0.95, 'area': 12000}, # car
 {'id': 3, 'label_id': 1, 'score': 0.92, 'area': 8500}, # another car
 ...
 ]
}
```

 
### Adding Segmentation to Pixeltable

 
> 
 
**Coming Soon:** Panoptic segmentation will be available as a built-in Pixeltable function. In the meantime, the UDFs below demonstrate how you can integrate any model yourself.

 

 
```python
import pixeltable as pxt
import PIL.Image
import numpy as np

@pxt.udf
def panoptic_segmentation(image: PIL.Image.Image) -> dict:
 """Run panoptic segmentation and return segment info."""
 import torch
 from transformers import DetrForSegmentation, DetrImageProcessor
 
 model_id = 'facebook/detr-resnet-50-panoptic'
 processor = DetrImageProcessor.from_pretrained(model_id)
 model = DetrForSegmentation.from_pretrained(model_id)
 
 with torch.no_grad():
 inputs = processor(images=image, return_tensors="pt")
 outputs = model(**inputs)
 
 result = processor.post_process_panoptic_segmentation(
 outputs, target_sizes=[(image.height, image.width)], threshold=0.85
 )[0]
 
 # Extract segment info (mask tensor is too large to store directly)
 segments = []
 for seg in result['segments_info']:
 mask = (result['segmentation'] == seg['id']).numpy()
 segments.append({
 'label': model.config.id2label.get(int(seg['label_id']), 'unknown'),
 'label_id': int(seg['label_id']),
 'score': float(seg.get('score', 1.0)),
 'area': int(mask.sum()),
 'is_thing': seg['label_id'] < 80, # COCO: 0-79 are things, 80+ are stuff
 })
 
 return {
 'num_segments': len(segments),
 'segments': segments,
 'things': [s for s in segments if s['is_thing']],
 'stuff': [s for s in segments if not s['is_thing']],
 }

# Add to frames view
frames.add_computed_column(segmentation=panoptic_segmentation(frames.frame))
```

 
### Visualizing Segmentation Masks

 
```python
@pxt.udf
def visualize_segmentation(image: PIL.Image.Image) -> PIL.Image.Image:
 """Create a colorized segmentation overlay."""
 import torch
 from transformers import DetrForSegmentation, DetrImageProcessor
 
 model_id = 'facebook/detr-resnet-50-panoptic'
 processor = DetrImageProcessor.from_pretrained(model_id)
 model = DetrForSegmentation.from_pretrained(model_id)
 
 with torch.no_grad():
 inputs = processor(images=image, return_tensors="pt")
 outputs = model(**inputs)
 
 result = processor.post_process_panoptic_segmentation(
 outputs, target_sizes=[(image.height, image.width)], threshold=0.85
 )[0]
 
 # Create color map
 seg_map = result['segmentation'].numpy()
 num_segments = len(result['segments_info'])
 
 # Generate distinct colors for each segment
 np.random.seed(42)
 colors = np.random.randint(0, 255, size=(num_segments + 1, 3), dtype=np.uint8)
 
 # Map segment IDs to colors
 colored = np.zeros((*seg_map.shape, 3), dtype=np.uint8)
 for i, seg in enumerate(result['segments_info']):
 mask = seg_map == seg['id']
 colored[mask] = colors[i]
 
 # Blend with original image
 overlay = PIL.Image.fromarray(colored)
 blended = PIL.Image.blend(image.convert('RGB'), overlay, alpha=0.5)
 
 return blended

# Add visualization column
frames.add_computed_column(seg_viz=visualize_segmentation(frames.frame))

# Create segmented video
frames.group_by(videos).select(
 pxt.functions.video.make_video(frames.pos, frames.seg_viz)
).show()
```

 
### Querying Segmentation Data

 
Once segmentation is stored, you can query it:

 
```python
# Find frames with many objects
frames.where(frames.segmentation.num_segments > 10).select(
 frames.frame, frames.segmentation.num_segments
).show()

# Count frames by dominant stuff category
frames.select(
 frames.segmentation.stuff[0].label # Most prominent stuff class
).show()

# Find frames containing specific objects
frames.where(
 frames.segmentation.things.apply(lambda t: any(s['label'] == 'car' for s in t))
).show()
```

 
### When is Segmentation Useful?

 
| Use Case | Bounding Boxes | Panoptic Segmentation |
| --- | --- | --- |
| Object counting | ✓ | ✓ |
| Object localization | ✓ | ✓ |
| Precise object boundaries | ✗ | ✓ |
| Scene composition analysis | ✗ | ✓ |
| Background removal | ✗ | ✓ |
| Image editing/compositing | ✗ | ✓ |
| Occlusion handling | ✗ | ✓ |
| Area/size estimation | Approximate | Precise |

 
### Use Bounding Boxes When:

 

 - You just need to know *where* objects are

 - Speed matters more than precision

 - You're doing object detection for downstream tasks

 

 
### Use Panoptic Segmentation When:

 

 - You need pixel-precise object boundaries

 - You're doing scene understanding or composition analysis

 - You need to separate foreground from background

 - You're building image editing or AR applications

 

 
### The Tradeoff

 
| Metric | Object Detection (DETR) | Panoptic Segmentation |
| --- | --- | --- |
| Inference time | ~100ms | ~200ms |
| Output size | ~1KB (boxes + labels) | ~1MB+ (full mask) |
| Use case | "What's in this image?" | "What's every pixel?" |

 
For most video analysis tasks, bounding boxes are sufficient. Segmentation adds value when you need precise boundaries or scene composition.

 
## Conclusion

 
Rerun and Pixeltable solve the same problem differently. Rerun optimizes for real-time streaming with sparse detection + tracking. Pixeltable optimizes for batch processing with detection on every frame.

 
The result? Pixeltable's [declarative approach](/blog/declarative-vs-imperative-ai-pipelines) eliminates the need for tracking entirely: what seems like a core feature of the Rerun example is actually just an optimization for streaming that becomes unnecessary in a different architecture.

 
Segmentation adds another dimension of understanding, but at a cost. Use it when you need pixel-level precision; skip it when bounding boxes suffice.

 
**Sometimes the best code is the code you don't have to write.**

 
## Related Resources

 

 - **[Pixeltable on GitHub](https://github.com/pixeltable/pixeltable)** - Get started with declarative video processing

 - **[Video Keyframe Extraction Guide](/blog/video-keyframe-extraction-pixeltable)** - Extract meaningful frames from videos

 - **[Declarative vs Imperative AI Pipelines](/blog/declarative-vs-imperative-ai-pipelines)** - Understand the paradigm shift

 - **[Python UDFs in Pixeltable](/blog/python-udfs-pixeltable)** - Create custom processing functions

 - **[Rerun.io](https://rerun.io)** - Explore Rerun's visualization capabilities

 - **[Join Pixeltable Discord](https://discord.gg/QPyqFYx2UN)** - Discuss video processing workflows