---
title: "Declarative, Multimodal, Incremental AI Infrastructure with Pixeltable"
date: "2025-02-25"
author: "Pierre Brunelle"
tags:
  - Declarative AI
  - Multimodal AI
  - Pixeltable
description: "Simplify AI data infrastructure with Pixeltable's declarative, multimodal, incremental approach. Free developers from pipeline complexity for video/image tasks."
url: "https://pixeltable.com/blog/declarative-multimodal-incremental"
---

# Declarative, Multimodal, Incremental AI Infrastructure with Pixeltable

## The Database Analogy: Focusing on What Matters

 
Think about how web developers interact with relational databases like Postgres or data warehouses like Snowflake. They don't typically worry about low-level file storage, B-tree indexing, transaction management protocols, or query execution planning. The database provides a simple, declarative table interface (SQL), freeing developers to focus on application logic and data modeling.

 
Why shouldn't AI development benefit from the same level of abstraction for its data infrastructure?

 
## Pixeltable's Approach: Declarative AI Infrastructure

 
Pixeltable applies this proven database philosophy to the unique challenges of AI workflows, particularly those involving complex multimodal data. It provides a familiar table interface, but one that seamlessly integrates both data storage *and* the transformation logic (feature extraction, model inference, etc.) applied to that data.

 
With Pixeltable, you control:

 

 - **What** transformations to apply (e.g., extract frames, detect objects, generate embeddings).

 - **Which** models or functions to use (built-in or custom).

 - **How** to structure your high-level algorithms and data flow using tables and views.

 

 
Pixeltable automatically handles the undifferentiated heavy lifting:

 

 - Workflow orchestration.

 - Transactional storage of data and metadata.

 - Efficient data retrieval and caching.

 - **Incremental updates:** Only recomputing results when necessary.

 - Data lineage and versioning.

 

 
The interface is designed for extensibility, allowing you to easily integrate your own Python functions (UDFs) into the declarative framework.

 
## Example: [Video Processing](/blog/video-analysis-guide) Simplified

 
This approach shines when dealing with notoriously complex data like video. Pixeltable includes built-in, optimized functionality for common tasks:

 

 - Efficiently referencing source video files.

 - Automatic frame extraction via iterators (e.g., `frame_iterator`).

 - Audio separation (e.g., `pxt.functions.video.extract_audio`).

 - Automated metadata extraction.

 

 
You can then extend this foundation by applying built-in or custom AI functions to the extracted frames or audio, such as object detection models, transcription services, or feature extractors.

 
For example, setting up frame extraction and object detection becomes remarkably concise:

 
```python

import pixeltable as pxt
from pixeltable.functions.video import frame_iterator
# Assume object detection function is available
# Example: from pixeltable.functions.vision import yolox

# 1. Create table referencing video files - Updated Type
videos = pxt.create_table('videos', {'video': pxt.Video})

# 2. Create view with automatic frame extraction (1 FPS)
# Pixeltable handles incremental processing
frames = pxt.create_view(
 'frames',
 videos,
 iterator=frame_iterator(video=videos.video, fps=1)
)

# 3. Apply object detection using add_computed_column - Updated Syntax
# Pixeltable manages execution, storage, lineage
frames.add_computed_column(detections=pxt.functions.vision.yolox(frames.frame))

# Now 'frames.detections' contains results, updated automatically
 
```

 
## Example: [Image Processing](/blog/video-similarity-search) Simplified

 
The same principles apply to image processing. Pixeltable handles:

 

 - Image loading and managed storage.

 - Basic transformations (available via functions or libraries like Pillow in UDFs).

 - Efficient batch processing during computation.

 

 
You can easily add steps like embedding generation for similarity search:

 
```python

import pixeltable as pxt
# Assume embedding function is available
# Example: from pixeltable.functions.huggingface import clip

# 1. Create an image table - Updated Type
images = pxt.create_table('images', {'image': pxt.Image})

# 2. Create an embedding index for fast similarity search - Updated Syntax
# Pixeltable manages index creation, updates, and embedding computation
# Example using CLIP model
images.add_embedding_index(
 'image', # The column to index
 embed=pxt.functions.huggingface.clip.using( # Pass the embedding function directly
 image=images.image,
 model_id='openai/clip-vit-base-patch32'
 )
)
# Ready for search: images.select(...).nearest(...)
 
```

 
## Key Benefits of the Declarative Approach

 

 - **Simplified Data Management:** No more complex scripts for extraction, format handling, or versioning. Define the structure, Pixeltable manages the data.

 - **Efficient Processing:** Incremental updates and intelligent caching minimize redundant computation, saving significant time and cost.

 - **[Declarative Interface](/blog/ai-functions-vs-pipelines):** Express complex multimodal pipelines as simple table/view operations and computed columns. Focus on *what* you want, not *how* to implement the plumbing.

 - **Developer Focus:** Maintain full control over your core algorithms and models (often in UDFs), while Pixeltable handles the complex, undifferentiated data infrastructure tasks. Spend more time innovating.

 

 
## Complete Workflow Example with UDF

 
Combine built-in functions with your custom logic seamlessly:

 
```python

import pixeltable as pxt
from pixeltable.functions.video import frame_iterator
# Assume object detection function is available
# Example: from pixeltable.functions.vision import yolox
import PIL.Image # Example dependency for UDF

# 1. Define your custom processing logic as a UDF
@pxt.udf
def custom_video_analytics(frame: PIL.Image.Image) -> dict:
 # Replace with your specialized analysis code
 # Example: calculate blurriness, detect specific scene type, etc.
 width, height = frame.size
 results = {'is_blurry': False, 'width': width, 'height': height}
 # ... (your analysis logic) ...
 return results

# 2. Manage data declaratively - Updated Type
videos = pxt.create_table('videos', {'video': pxt.Video})
frames = pxt.create_view(
 'frames',
 videos,
 iterator=frame_iterator(video=videos.video, fps=1)
)

# 3. Add computed columns using add_computed_column - Updated Syntax
frames.add_computed_column(detections=pxt.functions.vision.yolox(frames.frame)) # Object detection
frames.add_computed_column(analytics=custom_video_analytics(frames.frame)) # Your custom analysis

# Pixeltable now manages this entire pipeline:
# - Frame extraction and caching
# - Parallel execution of yolox and custom_video_analytics where possible
# - Incremental updates when new videos are added or functions change
# - Efficient storage and retrieval of all results
# - Lineage tracking for all computed data
 
```

 
## Conclusion: The Future is Declarative AI Infrastructure

 
Just as declarative interfaces revolutionized database interactions, Pixeltable brings this power to AI development. By handling the complexities of multimodal data management, incremental computation, and orchestration, Pixeltable allows you to focus on building sophisticated AI applications faster and more efficiently.

 
Stop wrestling with data pipelines and start leveraging the power of declarative AI infrastructure.

 

 - **[Build Your First Pixeltable Project](/blog/your-first-pixeltable-project)** - 10-minute hands-on tutorial

 - **[Try Pixeltable on GitHub](https://github.com/pixeltable/pixeltable)**

 - **[Read the Documentation](/docs/getting-started)**

 - **[Join our Discord Community](https://discord.gg/QPyqFYx2UN)**