---
title: "What We Learned Shipping Application Templates for AI Agents"
date: "2026-05-19"
author: "Pierre Brunelle"
tags:
  - Developer Experience
  - AI Agents
  - Templates
  - Starter Kit
  - LLM
  - Testing
  - Documentation
  - Production AI
description: "We shipped six application templates for Pixeltable, then had AI agents build and test them. Every bug class was a documentation gap. Every documentation gap was a pitfall the framework post predicted. Here is what broke, why it broke, and how the fix loop turns templates into the highest-leverage acquisition artifact in the stack."
url: "https://pixeltable.com/blog/shipping-application-templates-for-ai-agents"
---

# What We Learned Shipping Application Templates for AI Agents

**Summary:** The previous post laid out a [five-stage framework for designing software for LLM agents](/blog/designing-software-for-llms-as-customers). This post is the field report. We shipped six application templates for Pixeltable: multimodal RAG, video intelligence, agent, audio intelligence, content pipeline, and data lab. Then we had AI agents build and test them. Every bug class was a documentation gap. Every documentation gap was a pitfall the framework post predicted. Here is what broke, why it broke, and how the fix loop turns templates into the highest-leverage acquisition artifact in the stack.

 
## Templates Are the Missing Acquisition Layer

 
The original framework described a five-stage journey: Discovery, Acquisition, Surface Choice, Generation, Reinforcement. Our [Starter Kit](https://github.com/pixeltable/pixeltable-starter-kit) shipped three *structural patterns*: serving (`FastAPIRouter` plus `pxt service update`), backend (FastAPI API), and batch (pipeline script). These patterns answered "how do I wire Pixeltable to the outside world?" but left a harder question unanswered: **"What do I build with it?"**

 
Structural patterns are necessary but not sufficient. An agent asked to "build a RAG app" will not start from a blank `pyproject.toml` and a TOML route config. It will search for an existing RAG template, find a LangChain one on GitHub (15,000+ results), and write LangChain code. The agent never encounters your framework because it never gets past the intent-to-pattern mapping.

 
Application templates solve this by occupying the search slot directly. When an agent (or a human) searches for "video analysis pipeline," a `video-intel` template that scaffolds in one command is the intervention that prevents the [LangChain/pandas/FAISS default](/blog/pixeltable-vs-langchain-rag-comparison).

 
| What We Had | What Was Missing | What We Shipped |
| --- | --- | --- |
| 3 structural patterns (serving, backend, batch) | Domain-specific starting points | 6 application templates, each mapping to a pattern |
| uvx pixeltable-new myapp --serving | No way to scaffold a specific use case | uvx pixeltable-new --template multimodal-rag my-kb |
| SKILL.md with anti-patterns | Templates not referenced as entry points | Templates listed as first action in SKILL.md |

 
## The Six Templates

 
Each template is a complete, runnable project: schema, app/pipeline, dependencies, and (where appropriate) a web UI. Each maps to one or more structural patterns from the [Starter Kit](/blog/pixeltable-starter-kit-launch):

 
| Template | Pattern | What You Get |
| --- | --- | --- |
| multimodal-rag | serving + backend | Upload docs, images, video, audio; unified cross-modal search; LLM Q&A |
| video-intel | serving | Frame extraction, CLIP visual search, Whisper transcription, DETR object detection |
| agent | serving + backend | Tool-calling agent with persistent memory, knowledge base, conversation history |
| audio-intel | serving + backend | Audio upload, Whisper transcription, sentence-level search |
| content-pipeline | batch | Auto-detect media type, process images/docs/audio, export to Parquet/SQL |
| data-lab | batch | Image dataset management, CLIP search, DETR detection, PyTorch/COCO export |

 
The templates are scaffolded from the starter kit via the [`pixeltable-new`](https://github.com/pixeltable/pixeltable-new) CLI:

 
```bash
uvx pixeltable-new --template multimodal-rag my-kb
cd my-kb && uv sync
uv run python schema.py # create tables, views, indexes
uv run uvicorn app:app # start the API + UI
```

 
## The Testing Gauntlet: What Broke and Why

 
We tested every template end-to-end: scaffold, install, schema init, app start, HTTP endpoint verification. The results validated a core prediction from the [framework post](/blog/designing-software-for-llms-as-customers): **every runtime bug was traceable to a documentation gap, and every documentation gap was a pitfall that agents hit systematically.**

 
Five distinct bug classes emerged.

 
### Bug 1: Duplicate Embedding Indexes (4 of 6 templates)

 
```python
# The pattern that breaks
doc_chunks.add_embedding_index('text', string_embed=text_embed, if_exists='ignore')
```

 
When `schema.py` runs standalone and then gets re-imported by `app.py`, `add_embedding_index` is called twice. Without an explicit `idx_name`, each invocation generates an auto-name (`idx0`, `idx1`). The `if_exists='ignore'` check matches by name, does not find a duplicate, and creates a second index on the same column. The `@pxt.query` function then fails with `Column 'text' has multiple embedding indices; specify idx_name instead`.

 
**Fix:** Always specify `idx_name`:

 
```python
doc_chunks.add_embedding_index('text', idx_name='doc_text_idx',
 string_embed=text_embed, if_exists='ignore')
```

 
**Framework lesson:** This is a Generation-stage problem. The `if_exists='ignore'` idiom is documented and correct for tables and columns. But its behavior on embedding indexes has a subtle difference: name-based matching vs. column-based matching. No documentation mentioned this. The agent wrote code that *looked* idiomatic but was not. This is exactly the "hallucinate plausibly" failure mode from Stage 3.

 
### Bug 2: @pxt.query Results Used Imperatively (multimodal-rag)

 
```python
# Broken: treats @pxt.query return value as a DataFrame
results = search_documents(query_text, n=n).collect().to_pandas().to_dict('records')
```

 
`@pxt.query` functions return expression objects, not DataFrames. Calling `.collect()` on the result triggers `__getattr__`, which interprets `collect` as a JSON path access and raises a type error. The correct approach for imperative cross-modal search is direct table queries:

 
```python
dc = pxt.get_table('kb.doc_chunks')
sim = dc.text.similarity(string=query_text, idx='doc_text_idx')
results = dc.order_by(sim, asc=False).limit(n).select(dc.text, sim=sim).collect()
```

 
**Framework lesson:** This is the `@pxt.query` eager compilation pitfall. The function body is compiled at decoration time with expression placeholders, not executed as regular Python. The agent conflated two valid Pixeltable patterns (query functions for declarative routes and direct table queries for imperative code) into a broken hybrid. We had already documented this in the SKILL.md Common Pitfalls table as item #8, but the *template code itself* violated it.

 
### Bug 3: UDFs Defined in __main__ (video-intel)

 
```python
# Fails: Pixeltable requires UDFs to be in named modules
@pxt.udf
def _has_label(labels: list[str] | None, label: str) -> bool:
 return label in (labels or [])
```

 
Pixeltable cannot serialize UDFs defined in the global namespace of a `__main__` script. The UDF must live in a named module (e.g., `functions.py`). This constraint exists because Pixeltable persists UDF references for [incremental recomputation](/blog/economics-of-incremental-ai). It needs a stable module path to re-import the function.

 
**Fix:** Move UDFs to `functions.py`, import at the top of `schema.py`.

 
**Framework lesson:** The agent correctly identified that a UDF was the right tool, but placed it in the wrong location. This is a Surface Choice problem. The agent has no strong prior about Pixeltable's module serialization requirements because no other framework works this way.

 
### Bug 4: pxt.get_view() Does Not Exist (content-pipeline)

 
```python
# Broken: API does not exist
'doc_chunks': pxt.get_view('pipeline.doc_chunks').count()
# Fixed: views and tables use the same accessor
'doc_chunks': pxt.get_table('pipeline.doc_chunks').count()
```

 
The agent hallucinated a `pxt.get_view()` function by analogy with frameworks that distinguish between tables and views at the API level. In Pixeltable, views *are* tables from the query perspective. `pxt.get_table()` returns both.

 
**Framework lesson:** Classic hallucination from training priors. SQL distinguishes `SELECT * FROM view` from `SELECT * FROM table` syntactically, and most ORMs have separate accessors. The agent applied that prior. Our documentation said "views are created with `create_view`" but never explicitly said "retrieved with `get_table`." Implicit symmetry assumptions are where hallucinations breed.

 
### Bug 5: Thread-Unsafe Table References in FastAPI (multimodal-rag)

 
Pixeltable `Table` objects are bound to the thread that created them. Using a module-level table reference inside a FastAPI endpoint (which runs in a thread pool) causes `Table was accessed from a thread other than the one that constructed it`.

 
**Fix:** Call `pxt.get_table()` inside each endpoint function.

 
**Framework lesson:** This was already documented in SKILL.md as pitfall #10, but the template code and the official `deployment/overview.mdx` docs showed module-level table references. Agents reproduced the antipattern *even when SKILL.md said otherwise* because code examples in official docs override negative prompts in skill files. Fix the docs first.

 
## The Documentation Feedback Loop

 
The [framework post](/blog/designing-software-for-llms-as-customers) argued that Stage 4 (Reinforcement) creates a flywheel: measure, fix docs, re-measure. Our template testing validated this concretely. Every bug class produced a documentation improvement:

 
| Bug Class | Doc Updated | New Content |
| --- | --- | --- |
| Duplicate embedding indexes | SKILL.md, core-api.md | Always use explicit idx_name with add_embedding_index |
| @pxt.query eager compilation | SKILL.md Common Pitfalls #8 | Do not call .collect(), insert(), or reference uninitialized tables inside @pxt.query |
| UDFs in __main__ | Templates themselves | Always place @pxt.udf functions in a functions.py module |
| get_view() does not exist | core-api.md | Both tables and views use pxt.get_table() |
| Thread-unsafe table refs | SKILL.md #10, deployment/overview.mdx | Call pxt.get_table() per-request in FastAPI endpoints |

 
The Common Pitfalls table in our [SKILL.md](https://github.com/pixeltable/pixeltable-skill) grew from 7 entries to 10 during this process. Each entry is a **negative prompt**: the wrong code, the correct code, and a one-line explanation. This is exactly the Stage 3 lever the framework post described as the highest-leverage investment: negative guidance that *deflects* the prior rather than competing with it.

 
## Templates as Eval Fixtures

 
The template testing process is itself an eval. Each template is a self-contained user story ("build a multimodal RAG app" or "build a video analysis pipeline") with a concrete success criterion: `schema.py` initializes, the app starts, and all endpoints return 200.

 
This gives us something the framework post called for but had not yet shipped: **end-to-end functional evals that run in CI.** The [pixeltable-eval](https://github.com/pixeltable/pixeltable-eval) harness measures whether agents *write* correct code. Template testing measures whether the *reference code we ship* is correct. Both are necessary. The eval harness catches agent failures. Template testing catches documentation failures.

 
The testing protocol:

 
```bash
for template in multimodal-rag video-intel agent audio-intel content-pipeline data-lab; do
 uvx pixeltable-new --template $template /tmp/$template
 cd /tmp/$template && uv sync
 uv run python schema.py
 uv run uvicorn app:app --port $PORT # or: pxt service run app.py my_app --port $PORT
 curl -s -o /dev/null -w "%{http_code}" POST /api/search ...
 curl -s -o /dev/null -w "%{http_code}" GET /api/stats ...
done
```

 
## What This Means for the Framework

 
The five-stage framework holds up. But the template experience sharpened three claims.

 
**1. Templates are a Discovery artifact, not just Acquisition.** The original framework placed templates in Stage 2 (Acquisition). But templates also serve Stage 1 (Discovery) because they occupy search-result slots that would otherwise go to competitors. `uvx pixeltable-new --template multimodal-rag` is discoverable in a way that "read the SKILL.md and write a schema from scratch" is not.

 
**2. The code you ship is documentation.** Agents treat template code as ground truth. If the template uses `add_embedding_index` without `idx_name`, agents will reproduce that pattern everywhere. Templates are not just onboarding material. They are the strongest positive prompt in the stack, stronger than the SKILL.md because agents *copy* code, not prose.

 
**3. The documentation-is-wrong failure mode is worse than no documentation.** When our `deployment/overview.mdx` showed module-level table references in FastAPI endpoints, agents reproduced the antipattern *even when SKILL.md said otherwise*. Code examples in official docs override negative prompts in skill files. Fix the docs first.

 
## What We Are Shipping Next

 
| Intervention | Stage | Status |
| --- | --- | --- |
| Template CI: schema init + app start + endpoint smoke test on every PR | Reinforcement | In progress |
| idx_name required in all SKILL.md examples | Generation | Shipped |
| pxt.lint() rule for missing idx_name | Generation | Planned |
| Structured errors with fix_example for "multiple embedding indices" | Generation | Planned |
| --template flag in pixeltable-eval for template-specific evals | Reinforcement | Planned |

 
## The General Principle

 
**Your templates are your strongest positive prompt.** Agents copy code. If the code in your templates is wrong, your agents will be wrong at scale. If the code in your templates is right (right imports, right idioms, right error handling) agents will propagate those patterns into every application built on your framework.

 
Test your templates the way you test your SDK: automatically, on every commit, with functional assertions. The template is not a marketing artifact. It is a unit test for your documentation.

 
## Resources

 

 - [Pixeltable Starter Kit](https://github.com/pixeltable/pixeltable-starter-kit): 3 patterns + 6 application templates

 - [pixeltable-new](https://github.com/pixeltable/pixeltable-new): `uvx pixeltable-new --template multimodal-rag my-kb`

 - [Pixeltable Skill](https://github.com/pixeltable/pixeltable-skill): install with `npx skills add pixeltable/pixeltable-skill`

 - [Pixeltable Eval](https://github.com/pixeltable/pixeltable-eval): 16 evals measuring agent code quality

 - [Designing Software for LLMs as Customers](/blog/designing-software-for-llms-as-customers): the five-stage framework this post builds on

 - [Five Things Pixeltable Does That Competitors Cannot](/blog/five-things-pixeltable-does-competitors-cant): the compounding capabilities behind these templates

 - [Pixeltable Documentation](https://docs.pixeltable.com)

 - [Pixeltable on GitHub](https://github.com/pixeltable/pixeltable)