Pixeltable vs Neon

Neon is an exceptional serverless Postgres: instant database branching, scale-to-zero compute, and separation of storage and compute. Pixeltable is an AI data infrastructure engine: multimodal column types, declarative computed columns, automatic embedding index maintenance, and built-in API serving. Pick Neon when your core workload is relational Postgres with preview branches. Pick Pixeltable when your rows are audio, video, documents, and embeddings that must automatically transform on insert.

pip install 'pixeltable[serve]'
See how it works

Relational storage vs AI dataflow

SidePixeltableNeon
At a glance
  • Multimodal types (pxt.Image, pxt.Audio, pxt.Video, pxt.Document) validated and stored natively
  • Computed columns execute Whisper, CLIP, and sentence transformers in-engine on insert
  • EmbeddingIndex declared directly on table classes with automated incremental synchronization
  • FastAPIRouter declared alongside tables; zero glue code for production HTTP endpoints
  • Full PostgreSQL compatibility with standard SQL dialect, foreign keys, and extensions (pgvector)
  • Instant copy-on-write database branching for CI/CD, testing, and isolated staging environments
  • Scale-to-zero autoscaling compute saves costs during idle periods on development databases
  • Requires external Celery/Airflow workers, S3 storage, and manual backfill scripts for AI transforms

What actually differs

Neon excels at serverless Postgres elasticity, instant copy-on-write branching, and scale-to-zero pricing. Pixeltable excels when unstructured media requires automated transformation pipelines, incremental recomputation, and vector indexing in the schema.

FeaturePixeltableNeon
Core architecture
Application schema with DAG transformation engine & serving
Serverless relational PostgreSQL with separated storage & compute
Multimodal types
Native types (pxt.Audio, Video, Image, Document) with caching & validation
Raw bytea / text columns or external S3 URL references
AI pipeline orchestration
Built-in declarative computed columns; insert runs transformation DAG
External orchestrator required (Airflow, Celery, Temporal)
Embedding index maintenance
Declared on table class; updates incrementally and atomically on insert
Manual embedding generation and pgvector INSERT / UPDATE scripts
Model evolution & backfills
Swap model in schema; engine backfills only the delta in place
ALTER TABLE + custom batch Python backfill script + downtime risk
Database branching
Schema versioning & lineage tracking; no instant physical branch
Instant copy-on-write branching in seconds via storage engine
PostgreSQL compatibility
Postgres storage backend via SDK abstraction
Native Postgres 15/16; full SQL, foreign keys, triggers, and extensions
HTTP API serving
FastAPIRouter in the same Python file
Database only; requires standalone backend (FastAPI, Next.js, Express)
Free plan tier
Community tier: hosted compute + managed catalog + 50 GB media storage and 10 GB database storage
1 GB storage per project, 100 projects, 100 CU-hours/mo, scale-to-zero compute

Document RAG: Chunking, embeddings & vector index

Pixeltable: Views chunk documents and indexes update incrementally on write. Neon: Requires DDL, an external script to chunk and embed, and manual pgvector upsert logic.

Pixeltable

import pixeltable as pxt
from pixeltable.functions.document import document_splitter
from pixeltable.functions.huggingface import sentence_transformer
TableModel = pxt.model_base()
embed = sentence_transformer.using(
model_id='sentence-transformers/all-MiniLM-L6-v2'
)
class Docs(TableModel, name='docs'):
document: pxt.Document
title: pxt.String
class Chunks(
TableModel,
name='chunks',
base=Docs,
iterator=document_splitter(Docs.document, separators='sentence', limit=512),
):
__indexes__ = [pxt.EmbeddingIndex(text, embedding=embed)]
# Apply schema & insert: transforms and embeddings run automatically
# pxt schema update app.py rag
docs = pxt.get_table('rag.docs')
docs.insert([{'document': 'annual_report.pdf', 'title': '2025 Annual Report'}])
# Query vector similarity in-engine
chunks = pxt.get_table('rag.chunks')
sim = chunks.text.similarity(string='operating margin growth')
results = chunks.order_by(sim, asc=False).limit(5).select(chunks.text, chunks.title)

Neon

# 1. Database schema in Neon
# CREATE EXTENSION IF NOT EXISTS vector;
# CREATE TABLE docs (id UUID PRIMARY KEY, title TEXT, s3_url TEXT);
# CREATE TABLE doc_chunks (id UUID PRIMARY KEY, doc_id UUID, text TEXT, embedding vector(384));
import fitz # PyMuPDF
from sentence_transformers import SentenceTransformer
import psycopg2
from psycopg2.extras import execute_batch
model = SentenceTransformer('sentence-transformers/all-MiniLM-L6-v2')
conn = psycopg2.connect("postgresql://neondb_owner:[email protected]/neondb")
def ingest_document(doc_id, title, pdf_path):
doc = fitz.open(pdf_path)
chunks = [page.get_text() for page in doc]
embeddings = model.encode(chunks)
with conn.cursor() as cur:
cur.execute("INSERT INTO docs VALUES (%s, %s, %s)", (doc_id, title, pdf_path))
rows = [(str(uuid.uuid4()), doc_id, c, e.tolist()) for c, e in zip(chunks, embeddings)]
execute_batch(cur, "INSERT INTO doc_chunks VALUES (%s, %s, %s, %s)", rows)
conn.commit()
# Plus: Celery task runner, S3 download glue, retry on network failure...

Changing the embedding model on live data

When you upgrade your embedding model, Pixeltable backfills the delta automatically. Neon requires an ALTER TABLE, custom migration script, and pagination to avoid timeouts.

Pixeltable

# Update embedder in app.py:
new_embed = sentence_transformer.using(model_id='BAAI/bge-small-en-v1.5')
class Chunks(
TableModel,
name='chunks',
base=Docs,
iterator=document_splitter(Docs.document, separators='sentence', limit=512),
):
__indexes__ = [pxt.EmbeddingIndex(text, embedding=new_embed)]
# Run: pxt schema update app.py rag
# Pixeltable calculates the delta, computes new embeddings,
# and updates the index without duplicating rows or manual scripts.

Neon

# 1. Neon DDL migration
# ALTER TABLE doc_chunks ADD COLUMN new_embedding vector(384);
# 2. Write and run custom backfill script
new_model = SentenceTransformer('BAAI/bge-small-en-v1.5')
batch_size = 100
offset = 0
while True:
with conn.cursor() as cur:
cur.execute("SELECT id, text FROM doc_chunks WHERE new_embedding IS NULL LIMIT %s", (batch_size,))
rows = cur.fetchall()
if not rows:
break
ids, texts = zip(*rows)
vectors = new_model.encode(list(texts))
update_data = [(v.tolist(), i) for v, i in zip(vectors, ids)]
execute_batch(cur, "UPDATE doc_chunks SET new_embedding = %s WHERE id = %s", update_data)
conn.commit()
# Risk: transaction timeouts, connection drops, and drift between old and new columns.

When to choose which platform

Choose Pixeltable when

  • Multimodal AI is your core product

    Audio, video, PDFs, and embeddings. Pixeltable automates chunking, transcription, and indexing in one Python schema.

  • You want to eliminate pipeline glue code

    Replacing external orchestrators (Airflow, Celery) and backfill scripts with declarative computed columns.

  • You need Python-native ML integration

    Direct integration with PyTorch, Whisper, Hugging Face, OpenAI, and Anthropic without microservice hops.

Choose Neon when

  • Your core data is relational

    Standard transactional Postgres workloads with complex relational joins, foreign keys, and ACID guarantees.

  • You need preview environment database branches

    Neon instant copy-on-write database branching allows every pull request to have an isolated staging database.

  • You want scale-to-zero serverless pricing

    Development and preview databases spin down when idle, minimizing hosting costs for multi-tenant Postgres.

Making the right choice

  • Coexistence: Neon as Relational Store + Pixeltable as AI Dataflow

    • Many teams use Neon for primary transactional records (users, billing, core entities) and Pixeltable for AI workloads (multimodal media, embeddings, document intelligence).
    • Insert external Neon IDs into Pixeltable tables to correlate AI transformations with relational entities.
    • Eliminate custom Airflow/Celery jobs by letting Pixeltable manage the transformation DAG.

Frequently asked questions

The pipeline is the schema. Insert a row, transforms run.

Declare tables and computed columns in app.py. Apply with pxt schema update. No Airflow, Celery, or backfill scripts required.

pip install 'pixeltable[serve]'
See how it worksGet expert guidance