---
title: "Build Multimodal AI Applications Faster with Pixeltable Infrastructure"
date: "2025-03-10"
author: "Pierre Brunelle"
tags:
  - Multimodal AI
  - Pixeltable
  - AI Infrastructure
description: "Build complex multimodal AI applications (text, image, video, audio) faster with Pixeltable's unified declarative infrastructure. Simplify multimodal AI development and stop juggling multiple tools."
url: "https://pixeltable.com/blog/building-multimodal-apps"
---

# Build Multimodal AI Applications Faster with Pixeltable Infrastructure

## The Multimodal Challenge: Beyond Simple Data Types

 
Modern AI applications increasingly need to understand and process multiple types of data simultaneously – text, images, video clips, audio streams, PDFs, and more. Building systems that can seamlessly ingest, process, relate, and query across these modalities is complex. Traditional approaches often require stitching together numerous specialized tools, leading to brittle pipelines, data silos, and significant engineering overhead.

 
## The Pixeltable Solution: Unified & Declarative Infrastructure

 
Pixeltable provides a unified, [declarative data infrastructure](/blog/declarative-multimodal-incremental) designed specifically for the AI era. Instead of complex orchestration scripts, you define your data structures and processing logic using familiar table and view concepts, letting Pixeltable handle the underlying complexity of managing diverse data types, incremental updates, and data lineage.

 
## Example: Unified Table & Cross-Modal Processing

 
Imagine needing to process videos, extract audio, transcribe it, and then index the transcriptions for search. Here's how Pixeltable streamlines this:

 
```python

import pixeltable as pxt
# Assuming necessary functions are imported or registered
from pixeltable.functions.video import extract_audio
from pixeltable.functions.openai import transcriptions # Example, requires openai extra
from pixeltable.functions.string import string_splitter
from pixeltable.functions.sentence_transformer import encode as embed_text # Example embedding

# Assume client is initialized: client = pxt.Client()

# 1. Create a unified table for multimodal content
content_table = pxt.create_table('content_library', {
 'video': pxt.Video,
 'audio': pxt.Audio,
 'source_name': pxt.String
})

# 2. Audio extraction using add_computed_column
content_table.add_computed_column(extracted_audio=extract_audio(content_table.video))

# 3. Transcription using add_computed_column
# Uses extracted_audio if available, otherwise the original audio column
transcribe_target = content_table.extracted_audio.coalesce(content_table.audio)
content_table.add_computed_column(transcript=transcriptions(transcribe_target, model='whisper-1'))

# Now, inserting a video triggers audio extraction AND transcription automatically
# content_table.insert([{'video': 'path/to/my_video.mp4', 'source_name': 'Lecture 1'}])
 
```

 
This entire pipeline, from [video/audio ingestion](/blog/video-analysis-guide) to searchable transcript embeddings, is defined declaratively. Pixeltable manages the execution, dependencies, incremental updates, and lineage automatically.

 
## Example: Indexing & Search on Processed Data

 
To make the transcribed text searchable, we can create a view to split it into sentences and then add an embedding index:

 
```python

# 4. Filter the base table to get rows with valid transcripts
valid_transcripts = content_table.where(content_table.transcript['text'] != None)

# 5. Create a view based on the filtered data
# This view automatically updates based on the filtered rows
sentences_view = pxt.create_view(
 'transcription_sentences',
 valid_transcripts, # Base the view on the filtered result
 iterator=string_splitter(
 # Target the text part of the transcription result
 text=valid_transcripts.transcript['text'],
 separators='sentence' # Split by sentence
 )
)

# 6. Add embeddings using add_computed_column
sentences_view.add_computed_column(embedding=embed_text(sentences_view.text))

# 7. Create the vector index
sentences_view.add_embedding_index('embedding')

# Now you can perform semantic search on the video/audio transcripts!
# query_embedding = embed_text("search query")
# results = sentences_view.select(...).nearest(sentences_view.embedding, query=query_embedding)
 
```

 
## Key Feature: Unified Data Management

 
Stop juggling separate systems for different data types. Pixeltable provides a single, consistent table interface where video, audio, images, documents, and structured metadata coexist. This simplifies data loading, querying, and relationship management across modalities.

 
```python

# Example: Inserting different data types into one table schema
# (Assuming table schema includes nullable columns for each type)
multimodal_table = pxt.create_table('assets', {
 'asset_id': pxt.String,
 'image_data': pxt.Image,
 'video_data': pxt.Video,
 'text_content': pxt.String
})

multimodal_table.insert([
 {'asset_id': 'img_001', 'image_data': 'path/to/image.jpg'},
 {'asset_id': 'vid_001', 'video_data': 'path/to/video.mp4'},
 {'asset_id': 'txt_001', 'text_content': 'Some document text.'}
])
 
```

 
## Key Feature: Cross-Modal Processing & Search

 
Apply functions that bridge modalities. Extract audio from video, generate text descriptions from images, or search across different types using multimodal embedding models like CLIP.

 
```python

# Example: Finding images similar to a text description using CLIP
# Assume 'images' table has an embedding index created with CLIP
# from pixeltable.functions.huggingface import clip
# images.add_embedding_index(image_embed=clip(images.image))

text_query = "a photo of a dog playing fetch"

# The .nearest() call handles embedding the text query using the same CLIP model
similar_images = images.select(
 images.image_path,
 images.image # Select columns needed
).nearest(images.image, query=text_query, limit=5)

results = similar_images.collect()
 
```

 
## Key Feature: Advanced Search & RAG

 
Pixeltable makes building sophisticated [RAG context retrieval](/blog/production-rag-data-centric) straightforward. You can define reusable query functions that encapsulate your retrieval logic, which can then be applied automatically as computed columns.

 
Imagine a `messages_view` containing chat history with embeddings. You can define a query to find relevant context for new questions:

 
```python

# Assume messages_view has columns: message_id, username, text, embedding
# Assume messages_view has an embedding index on the 'embedding' column

# Define a reusable query function attached to the view
@messages_view.query
def get_context(question_embedding, threshold: float = 0.7, limit: int = 5):
 ""Finds relevant messages based on embedding similarity.""
 # Use .nearest() on the pre-indexed embedding column
 return messages_view.select(
 messages_view.text,
 messages_view.username,
 # Score is automatically included by .nearest()
 ).nearest(messages_view.embedding, query=question_embedding, limit=limit)
 # Optionally filter by similarity score threshold if needed *after* nearest
 # .where(messages_view.similarity > threshold)

# Now, imagine a 'questions' table:
# questions = pxt.create_table('questions', {'id': pxt.IntType(), 'question_text': pxt.String})
# questions['question_embedding'] = embed_text(questions.question_text)

# Apply the query function as a computed column:
# This automatically retrieves context for each question!
# questions['retrieved_context'] = get_context(questions.question_embedding)
 
```

 
This declarative approach for RAG context retrieval offers several advantages:

 

 - Retrieval logic is encapsulated and reusable.

 - Context retrieval runs automatically for new questions.

 - Results (the retrieved context) are stored and versioned alongside the questions.

 - The entire process is incremental and benefits from lineage tracking.

 

 
This is significantly cleaner and more maintainable than manual pipeline orchestration for RAG.

 
## Getting Started

 
Ready to start building? Try our [hands-on tutorial: Build a Smart Image Organizer in 10 Minutes](/blog/your-first-pixeltable-project) for a practical introduction, or create a basic multimodal table yourself:

 
```python

import pixeltable as pxt

# Define a table schema supporting multiple types
content = pxt.create_table('my_content', {
 'description': pxt.String,
 'main_video': pxt.Video,
 'thumbnail_image': pxt.Image,
 'tags': pxt.Json # Example for storing list of tags
})

# Insert data...
# content.insert([...])
 
```

 
## Conclusion: Focus on Innovation, Not Infrastructure

 
Building multimodal AI applications doesn't have to mean wrestling with a complex web of specialized tools and fragile integration code. Pixeltable provides the unified, declarative data infrastructure needed for the AI era, allowing you to focus on creating innovative applications while Pixeltable handles the underlying complexity.

 

 - **[Try Pixeltable on GitHub](https://github.com/pixeltable/pixeltable)**

 - **[Learn More About Multimodal Data](/docs/concepts/multimodal-data)**

 - **[Join our Discord Community](https://discord.gg/QPyqFYx2UN)**