---
title: "Kubrick Video Agent Course: Building Multimodal Agents with Pixeltable"
date: "2025-06-25"
author: "Pixeltable Team"
tags:
  - Video Agent
  - Multimodal AI
  - MCP
  - Model Context Protocol
  - Pixeltable
  - AI Course
  - Video Processing
  - Computer Vision
  - Agent Framework
description: "Learn how Pixeltable powers the new Kubrick video agent course - a hands-on deep dive into building production-ready multimodal agents for video processing and analysis."
url: "https://pixeltable.com/blog/kubrick-video-agent-course"
---

# Kubrick Video Agent Course: Building Multimodal Agents with Pixeltable

## The Future of Video Agents is Here

 
We're thrilled to highlight an exciting new hands-on course that showcases Pixeltable's power in building production-ready multimodal agents. [Miguel Otero Pedrido from The Neural Maze](https://theneuralmaze.substack.com/) and [Alex Razvant from Neural Bits](https://neuralbits.substack.com/) have launched **Kubrick** - a comprehensive course on building multimodal video agents using the Model Context Protocol (MCP).

 
> 
 
"**Every tool under the hood is powered by a library called Pixeltable.** Pixeltable handles incremental storage, transformation, indexing, and orchestration for multimodal data, which made it a perfect match for Kubrick, since we're working with video, images, and audio."

 Miguel Otero Pedrido, The Neural Maze
 

 
## Why Video Agents Matter

 
Video agents represent the next frontier in AI applications. Unlike traditional text-based agents, video agents can:

 

 - **Answer specific questions about video content** ("What's Morty's T-shirt color?")

 - **Clip scenes based on user queries** ("Show me the clip where HAL says 'I'm sorry Dave'")

 - **Find scenes using image similarity** (upload an image to find similar video frames)

 - **Process multimodal context** (video, audio, images, and metadata together)

 

 
## How Pixeltable Powers Kubrick

 
The course demonstrates why [Pixeltable's declarative multimodal infrastructure](/blog/unified-multimodal-ai-infrastructure-pixeltable) is ideal for video agent development:

 
### Declarative Video Processing

 
Instead of hand-coding complex video processing logic, Pixeltable's abstractions handle:

 

 - **Automatic frame extraction** from video files

 - **Audio transcription** with built-in AI functions

 - **Embedding generation** for semantic search

 - **Incremental updates** when new videos are added

 

 
### Seamless MCP Integration

 
The course shows how to build a production-ready MCP server that exposes Pixeltable's capabilities as standardized tools, making video processing accessible to any [MCP-compatible application](/blog/pixeltable-mcp-servers).

 
### Stateful Agent Architecture

 
Kubrick demonstrates using [Pixeltable as the persistence layer for stateful agents](/blog/building-memory-powered-ai-stateful-agents-pixeltable), enabling continuous learning and memory across conversations.

 
## Three-Component Architecture

 
The course breaks down Kubrick's architecture into three main components:

 
### 1. MCP Server for Video Processing

 
Built from scratch using [FastMCP](https://github.com/jlowin/fastmcp), with all tools powered by Pixeltable's multimodal capabilities.

 
### 2. Agentic API with MCP Clients

 
A FastAPI application that creates stateful agents, using Pixeltable for persistence and [Opik](https://github.com/comet-ml/opik) for observability.

 
### 3. HAL 9000-Inspired UI

 
A sleek interface that brings the video agent to life, complete with a video library for browsing processed content.

 
## Why This Course Matters for Pixeltable Users

 
This course is a perfect real-world example of [building production-ready AI agents](/blog/practical-guide-building-agents) with Pixeltable. It demonstrates:

 

 - How to leverage Pixeltable's [declarative approach](/blog/declarative-vs-imperative-ai-pipelines) for complex multimodal workflows

 - Integration patterns with modern AI frameworks and protocols

 - Best practices for building scalable, maintainable agent architectures

 - Real-world applications beyond simple demos

 

 
## Ready to Build Your Own Video Agent?

 
Whether you're interested in following the Kubrick course or building your own video agents, Pixeltable provides the foundation you need:

 
```python

import pixeltable as pxt
from pixeltable.functions.video import frame_iterator

# Create a table for videos
videos = pxt.create_table('videos', {
 'video': pxt.Video,
 'title': pxt.String
})

# Create a view that extracts frames
frames = pxt.create_view(
 'frames',
 videos,
 iterator=frame_iterator(video=videos.video, fps=1)
)

# Add AI-powered analysis
frames.add_computed_column(
 caption=openai.chat_completions(
 messages=[{
 'role': 'user',
 'content': [
 {'type': 'text', 'text': "Describe this frame"},
 {'type': 'image_url', 'image_url': {'url': frames.frame}},
 ],
 }],
 model='gpt-4o-mini',
 ).choices[0].message.content
)

# Add semantic search
frames.add_embedding_index('frame', image_embed=clip_image())
 
```

 
 
 Read the Full Course Introduction →
 
 

 
The complete course covers everything from MCP server development to custom observability layers. It's an excellent resource for anyone looking to understand how Pixeltable enables sophisticated multimodal AI applications.

 
## Additional Resources

 

 - **[Full Kubrick Course Introduction](https://theneuralmaze.substack.com/p/your-first-video-agent-multimodality)**

 - **[Pixeltable Video Processing Tutorial](/docs/tutorials/video-rag)**

 - **[Pixeltable MCP Servers Guide](/blog/pixeltable-mcp-servers)**

 - **[FastMCP - Fast, Pythonic MCP servers and clients](https://github.com/jlowin/fastmcp)**

 - **[Opik - Open-source LLM evaluation and observability](https://github.com/comet-ml/opik)**

 - **[Pixeltable GitHub Repository](https://github.com/pixeltable/pixeltable)**

 - **[Join our Discord Community](https://discord.gg/pixeltable)**