Pixeltable vs Supabase

Same video-intelligence app, two implementations. Pick Supabase when you want Postgres, row-level security, realtime, and a managed database. Pick Pixeltable when the pipeline is the product: video, frames, transcripts, embeddings, and retrieval in one Python file. They are not mutually exclusive.

pip install 'pixeltable[serve]'
See how it works

The trade

SidePixeltableSupabase
At a glance
  • 129 lines, one file, one process: ffmpeg, Whisper, and CLIP run in the schema
  • A row inserted by anything at all gets processed
  • Adding a title-embedding index is one line; the table backfills in place
  • Per-cell errors, lineage, and revert live in the catalog
  • 302 app lines plus 246 in compute-service, because Deno cannot run ffmpeg
  • RLS enabled on every table; one config line authenticates the Edge Function
  • Realtime, PITR, branching, and a managed database your team already operates
  • Fastest ingest at both tiers (16.89× realtime at 100 videos) and fastest evolve wall-clock at the larger corpus

What we measured

Same app. Supabase wins ingest at both tiers, RLS, realtime, and hosted-model latency. Pixeltable wins in-platform media, incremental schema, the local agent query, the evolve wall-clock at 23 videos, frame fetches at 203 videos, and the hosted swap on lines written.

FeaturePixeltableSupabase
Total code for this app
129 lines, 1 file
548 lines (302 + 246 compute-service), 7 files
ffmpeg, Whisper, CLIP
Computed columns, same process
compute-service; three of seven endpoints have no hosted-API substitute
Processing for any writer
Yes — the pipeline is the schema
No, unless you add database triggers
Add a derived column (lines)
1 line, 1 file, one command
24 lines, 2 files: ALTER TABLE plus a backfill script
Add a derived column (wall time)
0.96s at 23 videos, 5.22s at 123; schema change and backfill are the same step
2.71s at 23, 3.51s at 123; the winner flips with corpus
Ingest, 20 videos / 10 min
62.3s, 9.7× realtime
45.8s, 13.2× realtime
Ingest, 100 videos / 63 min
361.6s, 10.45× realtime
223.7s, 16.89× realtime
Frame search p50 (203 videos)
22.2ms
17.3ms
Agent query p50 (local model)
206ms; the model runs where the data is
700ms; three compute-service round trips per question
Reads at 203 videos, p50
5.1ms list, 0.7ms frame fetch
7.8ms list, 2.1ms frame fetch; every request traverses Kong
Search under 8 clients, p50
119.5ms; the in-process embedding serializes
67.0ms; the stateless layer stays parallel
Hosted model, 12 agent questions (paid endpoint)
12/12; 8.4s p50, the request-rate scheduler serializes pacing
12/12; 2.3s p50; zero recorded retries
Hosted swap, lines of retry code written
7; the provider scheduler paces and retries
36; pacing, Retry-After and backoff by hand
Authenticated endpoints
Open in this repo
One line: withSupabase({ auth: 'secret' })
Row-level security
None here
Enabled and verified on all five tables
Realtime push
None here
Built in
Per-cell errors and lineage
errormsg / errortype; pxt dashboard draws what produced a column
A failed step leaves NULL; Studio does not record lineage
Vendor checker in CI
ruff, generic Python; no Pixeltable conformance checker
deno lint and supabase db advisors on a live database

All three beat realtime on this laptop

Two corpus tiers: 20 videos ingested, then 100 more for 203 total. CPU, local models, not Cloud. Bold is best on that row.

Ingest

FeaturePixeltableSupabase
Wall time · 20 videos62.3s45.8s
Wall time · 100 videos361.6s223.7s
Faster than realtime · 20 videos9.7x13.18x
Faster than realtime · 100 videos10.45x16.89x
Median video · 20 videos2.9s2.0s
Median video · 100 videos3.5s2.2s

Search

FeaturePixeltableSupabase
Frame search p50 · 23 videos20.9ms20.4ms
Frame search p50 · 203 videos22.2ms17.3ms
Frame search p95 · 23 videos22.4ms26.5ms
Frame search p95 · 203 videos24.2ms21.2ms
Transcript search p50 · 23 videos17.9ms17.0ms
Transcript search p50 · 203 videos17.0ms16.6ms
Transcript search p95 · 23 videos21.1ms21.9ms
Transcript search p95 · 203 videos18.5ms22.2ms

Agent

Local Qwen2.5-1.5B on the 3-video baseline; retrieval is a fixed top-4 per index, so generation dominates.

FeaturePixeltableSupabase
Agent query p50206.1ms700.3ms
Agent query p95235.0ms851.1ms

Reads

FeaturePixeltableSupabase
GET /videos p50 · 23 videos6.1ms6.4ms
GET /videos p50 · 203 videos5.1ms7.8ms
Frame fetch p50 · 23 videos1.7ms2.8ms
Frame fetch p50 · 203 videos0.7ms2.1ms

Under load, 8 clients

The same ten searches with eight clients in flight; measures degradation, not speed.

FeaturePixeltableSupabase
Concurrent p50 · 23 videos173.6ms64.9ms
Concurrent p50 · 203 videos119.5ms67.0ms
Concurrent p95 · 23 videos264.6ms197.3ms
Concurrent p95 · 203 videos173.7ms123.9ms

Large tier ingests 20 videos (603.7s of footage) into a 3-video table; xl ingests 100 more (3,778s) for a 203-video, 7,689-frame corpus. All three finished every video at both tiers; Pixeltable logged one retried attempt at xl. Supabase led ingest at both tiers and the gap widened rather than shrank: 1.36x at large, 1.62x at xl. Search separates by single-digit milliseconds, because most of every number is embedding the query, not the index. The agent runs on the 3-video baseline, where Pixeltable is 3-4x faster because one question costs the other two three compute-service round trips and costs it none. Under eight concurrent search clients that same property inverts: the in-process embedding serializes at 120-174ms while the other two hold ~65ms. One machine, one afternoon: read the gaps, not the milliseconds.

One hosted model, three ways to call it

The agent again, with local generation swapped for nvidia/nemotron-3-super-120b-a12b on OpenRouter’s paid endpoint, identical for all three. 12 questions, 6 workers, then every patch reverts.

FeaturePixeltableSupabase
Answered of 121212
Wall time36.6s9.4s
p508.4s2.3s
p9534.1s5.2s
Retriesscheduler-internal0
Lines written for the swap736

The swap is the measurement: on Pixeltable it is a 7-line schema change and the request-rate scheduler paces and retries; on the other two it is a 36-37 line helper, because pacing, Retry-After and backoff are application code. On a healthy pool that code is never exercised: all three answer every question, both loops record zero retries, so the remaining differences are who wrote the retry code and latency, where provider response time dominates and the longest tail stays on Pixeltable, whose scheduler serializes pacing. The failure that code exists to catch is real, and now measured rather than inferred: probing the same model’s free pool returned HTTP 200 carrying an upstream error and no content on 3 of 36 calls. A hand-written loop inspects the body and retries that shape; a computed column evaluates the response it is given, so a malformed 200 lands as a null answer, since the scheduler retries raised errors, not well-formed wrong ones. An earlier free-pool run recorded exactly that: five empty cells, 7/12 answers where both loops held 12/12.

Adding a column to live data

One computed title-embedding index on a populated catalog. Lines compound; these seconds do not.

FeaturePixeltableSupabase
Total · 23-video catalog0.96s2.71s
Total · 123-video catalog5.22s3.51s
Schema change · 23 videos0.96s0.08s
Schema change · 123 videos5.22s0.09s
Model-free control · 23 videos0.53s0.23s
Model-free control · 123 videos0.81s0.12s
Backfill · 23 videossame step2.63s
Backfill · 123 videossame step3.42s
Lines written124
Files touched12
  • The wall-clock winner flips with corpus: Pixeltable is cheapest at 23 videos (0.96s), Supabase at 123 (3.51s). The fused step scales with rows while the script’s fixed cost amortises; a model-free control per platform separates mechanism cost from backfill work.
  • After pxt schema update, an insert against the already-registered route answers 409 until pxt service update. Reads keep working. The other two resolve the table on every request.
  • Backfill time is the part that scales, and this corpus cannot show it. Pixeltable’s backfill is work proportional to the rows that changed; a backfill script is work proportional to the table.

Ingest a video

Insert a video. Frames, audio, transcripts, embeddings, and scenes have to exist after that. On Pixeltable they are the schema. On the other two they live in the ingest path and in a second service.

Pixeltable

class Videos(TableModel, name='videos'):
video: pxt.Video
title: pxt.String
audio = extract_audio(video, format='mp3')
duration_sec = pxtf.video.get_duration(video)
scenes = video.scene_detect_content(threshold=8.0)
class Frames(TableModel, name='frames', base=Videos,
iterator=frame_iterator(Videos.video, fps=1.0)):
still = pxtf.image.resize(frame, (320, 180))
__indexes__ = [pxt.EmbeddingIndex(frame, embedding=VISUAL)]
class Chunks(TableModel, name='chunks', base=Videos,
iterator=audio_splitter(Videos.audio, duration=10.0)):
transcript = transcribe(audio_segment, model='base.en').text.astype(pxt.String)
__indexes__ = [pxt.EmbeddingIndex(transcript, embedding=SEMANTIC)]
Videos.insert([{'video': 'lecture.mp4', 'title': 'CS101'}])

Supabase

const { frames } = await compute("/extract-frames", { video_url, fps: FRAME_FPS });
const { embeddings } = await compute("/embed-clip", { images_b64: frames });
const frameRows = await Promise.all(frames.map(async (b64, i) => {
const path = `videos/${videoId}/frame_${i}.jpg`;
await supabase.storage.from("frames").upload(path, decodeBase64(b64), {
contentType: "image/jpeg", upsert: true,
});
return { video_id: videoId, frame_idx: i, embedding: embeddings[i] };
}));
await supabase.from("frames").insert(frameRows);

Search frames

Find frames of a whiteboard. Pixeltable asks the index. The other two embed the query themselves, then join or fetch rows in a second step.

Pixeltable

sim = Frames.frame.similarity(string=query)
return (
Frames.order_by(sim, asc=False)
.limit(limit)
.select(
frame_url=Frames.still,
frame_idx=Frames.pos,
video_title=Frames.title,
similarity=sim,
)
)

Supabase

CREATE FUNCTION search_frames(query_embedding vector(512), match_count INT)
RETURNS TABLE(frame_url TEXT, frame_idx INT, video_title TEXT, similarity FLOAT) AS $$
SELECT f.frame_url, f.frame_idx, v.title,
1 - (f.embedding OPERATOR(public.<=>) query_embedding)
FROM public.frames f
JOIN public.videos v ON f.video_id = v.id
ORDER BY f.embedding OPERATOR(public.<=>) query_embedding
LIMIT match_count;
$$ LANGUAGE sql STABLE;

Add a column to live data

Make the video title semantically searchable. The rows already exist. An embedding is not derivable in SQL, so every existing row has to be read, sent to a model, and written back.

Pixeltable

class Videos(TableModel, name='videos'):
...
__indexes__ = [pxt.EmbeddingIndex(title, embedding=SEMANTIC)]
# pxt schema update app.py media
# updated media/videos
# unchanged media/frames, media/chunks, media/conversations

Supabase

ALTER TABLE videos ADD COLUMN IF NOT EXISTS title_embedding vector(384);
CREATE INDEX IF NOT EXISTS videos_title_embedding_idx ON videos
USING hnsw (title_embedding vector_cosine_ops);
// then a script, because Postgres cannot call a model:
const { data: rows } = await db.from("videos")
.select("id,title").is("title_embedding", null);
for (let i = 0; i < rows.length; i += BATCH) {
/* embed the batch, update each row */
}

When to choose which platform

Choose Pixeltable when

  • The pipeline is the product

    Video, audio, images, documents, embeddings, and a retrieval step over them. The whole backend is one file, and nothing extra has to exist to run ffmpeg.

  • Schema changes have to stay cheap

    Adding a column backfills only that column. A row inserted by anything at all gets processed. That compounds; a line-count difference does not.

  • You already have an app backend

    Pixeltable as the media and retrieval layer behind a Supabase application is a coherent architecture, and for a team that already runs Postgres it is likely cheaper than moving.

Choose Supabase when

  • You want Postgres and the things around it

    Realtime subscriptions, row-level security for multi-tenancy, an auto-generated REST API, PITR, database branching, or a managed database your team already knows how to operate. Media work will live in a second service. That is the trade.

  • Throughput at this scale is the binding constraint

    On 20 videos and 10 minutes of footage, Supabase ingests fastest of the three, and again at 100 videos. If that is the constraint, follow it rather than the sponsor.

Making the right choice

  • What this page does not measure

    • Cost in dollars, multi-tenant authorization against auth.uid(), p99 latency, and on-call.
    • Reproduce the numbers: https://github.com/pixeltable/pixeltable-vs-supabase-vs-convex

Frequently asked questions

One file. The whole pipeline.

Declare the tables. Apply the schema. Insert a row. Serve the same file.

pip install 'pixeltable[serve]'
See how it worksGet expert guidance