---
title: "Building AI Data Infrastructure: Inside Pixeltable's Development Architecture"
date: "2025-01-21"
author: "Pixeltable Team"
tags:
  - AI Infrastructure
  - Architecture
  - Development
  - Multimodal Data
  - Python
description: "Explore the architectural decisions, development patterns, and engineering principles behind Pixeltable's declarative AI data infrastructure. Learn how we built a platform that handles multimodal data processing with incremental computation and type safety."
url: "https://pixeltable.com/blog/building-ai-data-infrastructure-pixeltable-architecture"
---

# Building AI Data Infrastructure: Inside Pixeltable's Development Architecture

Building production-ready AI applications requires more than just model integration. It demands a robust data infrastructure that can handle the complexity of multimodal data, ensure incremental processing, and maintain type safety throughout the entire pipeline. At Pixeltable, we've architected a declarative AI data infrastructure that addresses these challenges head-on.

 
In this deep dive, we'll explore the architectural decisions, development patterns, and engineering principles that make Pixeltable a powerful foundation for AI application development. Whether you're building similar systems or looking to understand modern AI infrastructure design, this post will give you insights into how we've tackled some of the most complex challenges in the space.

 
## The Declarative Philosophy

 
Traditional AI workflows are often imperative: you write scripts that explicitly define every step of data transformation, model inference, and result storage. This approach breaks down quickly as applications scale, leading to complex dependency management, redundant computation, and brittle pipelines.

 
Pixeltable takes a fundamentally different approach: **declarative data processing**. Instead of writing scripts, developers define what they want computed, and our engine determines how to execute it efficiently.

 
```python
# Traditional imperative approach
def process_videos(video_paths):
 results = []
 for path in video_paths:
 frames = extract_frames(path, fps=1.0)
 embeddings = []
 for frame in frames:
 embedding = clip_model.encode(frame)
 embeddings.append(embedding)
 results.append({
 'video_path': path,
 'frame_count': len(frames),
 'embeddings': embeddings
 })
 return results

# Pixeltable declarative approach
frames = pxt.create_view(
 'video_frames', videos,
 iterator=frame_iterator(video=videos.video, fps=1.0)
)
frames.add_computed_column(
 embedding=clip.using(model_id='openai/clip-vit-base-patch32')(frames.frame)
)
```

 
This declarative model brings several key advantages:

 
 

 - **Incremental Updates**: Only process new or changed data

 - **Automatic Dependency Tracking**: Changes propagate through the pipeline automatically

 - **Optimized Execution**: The engine can batch operations, push computations to SQL, and parallelize where appropriate

 - **Reproducible Results**: Same declarations always produce the same results

 

 
## Core Architectural Components

 
Pixeltable's architecture is built around several key components that work together to provide a unified interface for multimodal data processing.

 
### Unified Table Interface

 
At the heart of Pixeltable is a table abstraction that handles both structured and unstructured multimodal data. Unlike traditional databases that struggle with complex data types, our tables natively support images, videos, audio, and documents alongside standard data types.

 
```python
# Create table with multimodal schema
media_table = pxt.create_table('media_analysis', {
 'id': pxt.Int,
 'video': pxt.Video, # Local path or URL
 'thumbnail': pxt.Image, # PIL.Image.Image in memory
 'transcript': pxt.String, # Standard string
 'metadata': pxt.Json, # JSON data
 'embeddings': pxt.Array # NumPy arrays
})
```

 
The type system is built on specialized column types that understand the semantics of different media formats:

 

 - `pxt.Image`: PIL.Image.Image objects with file URL storage

 - `pxt.Video`: Video files with frame extraction capabilities

 - `pxt.Audio`: Audio data with waveform processing

 - `pxt.Document`: Text documents with parsing support

 - `pxt.Array`: NumPy arrays for embeddings and vectors

 

 
### Expression System

 
Pixeltable's expression system provides SQL-like operations that work seamlessly with multimodal data. Expressions can be composed, chained, and optimized by our execution engine.

 
```python
class MediaExpr(Expr):
 def __init__(self, operand: Expr, param: int):
 super().__init__(operand.col_type)
 self.operand = operand
 self.param = param
 self.components = [operand] # Dependencies
 
 def sql_expr(self) -> sql.ClauseElement | None:
 ""Return SQL if expressible in SQL, None otherwise.""
 if self.operand.sql_expr() is not None:
 return sql.func.media_function(self.operand.sql_expr(), self.param)
 return None
 
 def eval(self, data_row: DataRow, row_builder: RowBuilder) -> Any:
 ""Evaluate in Python if not expressible in SQL.""
 operand_val = data_row[self.operand.slot_idx]
 return process_media(operand_val, self.param)
```

 
The expression system intelligently decides whether to push computation to SQL or evaluate in Python, enabling optimal performance across different operation types.

 
### User-Defined Function (UDF) Framework

 
The UDF framework is where Pixeltable's extensibility really shines. Developers can create custom functions that integrate seamlessly with the declarative system while maintaining type safety and performance.

 
```python
# Simple UDF with type hints (required)
@pxt.udf
def extract_features(image: PIL.Image.Image, threshold: float = 0.5) -> dict[str, Any]:
 ""Extract image features with specified threshold.""
 # Implementation here
 return {'feature_count': 42, 'avg_intensity': 128.5}

# Batched UDF for efficiency
@pxt.udf(batch_size=32)
def batch_classify(images: Batch[PIL.Image.Image]) -> Batch[dict]:
 ""Classify multiple images efficiently.""
 # Process batch of 32 images at once
 return [{'class': 'cat', 'confidence': 0.95} for img in images]

# Async UDF for I/O operations
@pxt.udf
async def analyze_with_api(prompt: str) -> dict:
 ""Call external API for analysis.""
 async with httpx.AsyncClient() as client:
 response = await client.post('/analyze', json={'prompt': prompt})
 return response.json()
```

 
Our UDF framework enforces several key principles:

 

 - **Type Safety**: All UDFs must include type hints and return types

 - **Batching Support**: Automatic batching for expensive operations

 - **Resource Management**: Built-in rate limiting and connection pooling

 - **Error Handling**: Graceful failure handling with context preservation

 

 
### Execution Engine

 
The execution engine is responsible for translating declarative specifications into optimized execution plans. It handles incremental computation, dependency tracking, and performance optimization.

 
Key features of the execution engine include:

 

 - **Incremental Processing**: Only compute what's changed since last execution

 - **SQL Pushdown**: Move computations to the database when possible

 - **Parallel Execution**: Automatic parallelization of independent operations

 - **Resource Pooling**: Intelligent management of API rate limits and connections

 

 
## Development Best Practices

 
Building a robust AI infrastructure requires strict development practices. Here are the key principles we follow at Pixeltable.

 
### Type Safety First

 
Every piece of Python code in Pixeltable includes comprehensive type hints. This isn't just for documentation. It's enforced by our CI/CD pipeline and enables powerful static analysis.

 
```python
# Required: Type hints for all functions
def process_media_batch(
 inputs: list[PIL.Image.Image], 
 model_config: ModelConfig,
 batch_size: int = 32
) -> list[ProcessingResult]:
 ""Process batch of media with specified configuration.
 
 Args:
 inputs: List of PIL images to process
 model_config: Configuration for the processing model
 batch_size: Number of items to process in each batch
 
 Returns:
 List of processing results with confidence scores
 ""
 results: list[ProcessingResult] = []
 
 for i in range(0, len(inputs), batch_size):
 batch = inputs[i:i + batch_size]
 batch_results = model_config.process_batch(batch)
 results.extend(batch_results)
 
 return results
```

 
### Comprehensive Testing Strategy

 
Our testing strategy includes multiple levels of verification:

 
```bash
# Core testing commands
make test # Run pytest, stresstest, and quality checks
make fulltest # Comprehensive testing including expensive tests
make pytest # Unit tests only
make stresstest # Random table operations and stress testing
make nbtest # Test all Jupyter notebooks
```

 
Tests are designed with several key principles:

 

 - **Parallel Execution**: Tests run in parallel with automatic worker database assignment

 - **Automatic Retry**: Flaky tests are retried automatically to handle timing issues

 - **Isolation**: Each test worker gets its own database to prevent interference

 - **Comprehensive Coverage**: Unit tests, integration tests, notebook tests, and stress tests

 

 
### AI/ML Integration Patterns

 
We've developed consistent patterns for integrating AI/ML services that ensure reliability, performance, and maintainability:

 
```python
# Client registration pattern
@env.register_client('openai')
def _(api_key: str, base_url: str | None = None) -> 'openai.Client':
 return openai.Client(api_key=api_key, base_url=base_url)

# Rate-limited UDF with resource pool
@pxt.udf(resource_pool='request-rate:openai:chat')
async def chat_completions(
 messages: list[dict[str, str]], 
 *, 
 model: str, 
 model_kwargs: dict[str, Any] | None = None
) -> dict:
 ""OpenAI chat completions with automatic rate limiting.""
 client = env.Env.get().get_client('openai')
 result = await client.chat.completions.create(
 messages=messages, 
 model=model, 
 **(model_kwargs or {})
 )
 return result.model_dump()
```

 
This pattern ensures:

 

 - **Automatic Rate Limiting**: Resource pools handle API quotas

 - **Connection Reuse**: Clients are managed centrally

 - **Error Handling**: Consistent error propagation and retry logic

 - **Configuration Management**: Environment-based client configuration

 

 
## Performance & Scalability

 
Performance is critical for AI workloads. Pixeltable employs several strategies to ensure optimal performance across different scales and use cases.

 
### Intelligent Batching

 
Many AI operations benefit from batching, but the optimal batch size varies by operation type, hardware, and data characteristics. Pixeltable automatically handles batching based on UDF declarations:

 
```python
# Automatic batching for expensive operations
@pxt.udf(batch_size=16) # Process 16 images at once
def image_classification(images: Batch[PIL.Image.Image]) -> Batch[dict]:
 ""Classify images using a vision model.""
 # GPU operations are more efficient with larger batches
 model = load_classification_model()
 return model.predict_batch(images)

# Variable batch size based on input size
@pxt.udf(batch_size=lambda inputs: min(32, len(inputs)))
def adaptive_processing(texts: Batch[str]) -> Batch[dict]:
 ""Process texts with adaptive batching.""
 return process_text_batch(texts)
```

 
### SQL Pushdown

 
When possible, Pixeltable pushes computations to the PostgreSQL database, leveraging its optimized query engine:

 
```python
class FilterExpr(Expr):
 def sql_expr(self) -> sql.ClauseElement | None:
 ""Push simple filters to SQL when possible.""
 if isinstance(self.condition, SimpleComparison):
 return sql.and_(
 self.operand.sql_expr() > self.threshold,
 self.operand.sql_expr().isnot(None)
 )
 return None # Fall back to Python evaluation
```

 
### Resource Pooling

 
AI services often have rate limits and connection constraints. Pixeltable's resource pooling system manages these automatically:

 
```python
# Different pool types for different constraints
@pxt.udf(resource_pool='request-rate:openai:embeddings') # API rate limiting
@pxt.udf(resource_pool='gpu:cuda:0') # GPU resource management 
@pxt.udf(resource_pool='memory:high') # Memory-intensive operations
```

 
## Real-World Applications

 
The architectural decisions we've made enable powerful real-world applications. Here are some examples of how these patterns come together:

 
### Multimodal RAG System

 
```python
# Create document table with multimodal content
docs = pxt.create_table('documents', {
 'title': pxt.String,
 'pdf_file': pxt.Document,
 'images': pxt.Array[pxt.Image] # Extracted images
})

# Create processing pipeline
docs.add_computed_column(
 text_content=extract_text(docs.pdf_file)
)
docs.add_computed_column(
 text_embedding=openai.embeddings(docs.text_content, model='text-embedding-ada-002')
)

# Add embedding index for similarity search
docs.add_embedding_index('text_content', embedding=docs.text_embedding)

# Query with natural language
results = docs.where(docs.text_content.similarity(string='machine learning') > 0.8)
```

 
### Video Analysis Pipeline

 
```python
# Video processing with frame extraction
videos = pxt.create_table('videos', {'video_file': pxt.Video})

# Create frame view with automatic iteration
frames = pxt.create_view(
 'video_frames', videos,
 iterator=frame_iterator(video=videos.video_file, fps=1.0)
)

# Add computer vision processing
frames.add_computed_column(
 objects=yolo_detection(frames.frame)
)
frames.add_computed_column(
 scene_description=blip_captioning(frames.frame)
)

# Aggregate results back to video level
videos.add_computed_column(
 total_objects=frames.group_by(frames.video_id).objects.count()
)
```

 
## Getting Started with Development

 
Setting up a development environment for Pixeltable follows our established patterns for reproducibility and developer experience:

 
```bash
# Environment setup (Python 3.10 required)
conda create --name pxt python=3.10
conda activate pxt

# Install development environment
make install # Installs uv, ffmpeg, dependencies, jupyter kernel

# Development workflow
make test # Run tests before making changes
# ... make your changes ...
make check # Type checking, linting, formatting
make test # Run tests again
```

 
### Essential Development Commands

 
```bash
# Code quality
make check # Run all quality checks
make typecheck # MyPy type checking
make lint # Ruff linting 
make format # Ruff code formatting

# Testing
make test # Standard test suite
make fulltest # Comprehensive testing
make stresstest # Stress and chaos testing

# Documentation
mkdocs serve # Local documentation server
make release-docs # Deploy API documentation
```

 
## Future Architecture Directions

 
As AI infrastructure continues to evolve, we're exploring several architectural enhancements:

 

 - **Distributed Execution**: Scaling beyond single-machine processing

 - **Advanced Optimization**: Query planning and cost-based optimization

 - **Multi-Cloud Support**: Seamless integration across cloud providers

 - **Real-time Processing**: Streaming data support for live applications

 

 
## Building the Future of AI Infrastructure

 
The architectural decisions behind Pixeltable reflect our core belief: AI infrastructure should be **declarative**, **type-safe**, and **optimized for developer productivity**. By treating data transformations as first-class citizens and providing a unified interface for multimodal data, we've created a foundation that scales from prototypes to production.

 
The patterns and practices we've shared (from declarative processing to comprehensive testing) represent hard-learned lessons from building production AI systems. As the AI landscape continues to evolve, these architectural principles provide a stable foundation for innovation.

 
Whether you're building your own AI infrastructure or evaluating existing solutions, we hope these insights help you make informed decisions about architecture, development practices, and performance optimization. The future of AI applications depends on robust, scalable infrastructure, and we're excited to be pushing the boundaries of what's possible.

 
> 
 
Ready to experience declarative AI infrastructure firsthand? [Explore Pixeltable on GitHub](https://github.com/pixeltable/pixeltable) and see how our architectural decisions translate into developer productivity and application performance.