Pixeltable vs Convex

Same video-intelligence app, two implementations. Pick Convex when reactivity is the point — the UI re-renders on write. Pick Pixeltable when the pipeline is the product. A REST-shaped benchmark is Convex’s worst event; discount this column accordingly.

pip install 'pixeltable[serve]'
See how it works

The trade

SidePixeltableConvex
At a glance
  • 129 lines, one file: the pipeline is the schema
  • ffmpeg, Whisper, and CLIP run in-process; no second service
  • Declared HTTP routes; a missing query is a 422 before any handler runs
  • Adding a column backfills in place; processing fires for any writer
  • 425 app lines plus 246 in compute-service, and the easiest install of the three
  • npx convex dev: anonymous local backend, no account, no Docker
  • 105 lines in videos.ts because an action cannot write to the database
  • 90 lines in http.ts because this contract asked for REST instead of the reactive client

What we measured

Same app. Convex wins install, transcript search, reads, concurrent load, and hosted-model latency. Pixeltable wins in-platform media, the local agent query, and the hosted swap on lines written. REST is Convex’s worst event.

FeaturePixeltableConvex
Local install
pip install; large Python deps (torch, whisper, sentence-transformers)
npx convex dev — no account, no Docker, easiest of the three
Total code for this app
129 lines, 1 file
671 lines (425 + 246 compute-service), 7 files
ffmpeg, Whisper, CLIP
Computed columns, same process
compute-service; the Convex runtime cannot run them
Writes from an action
Insert is a row; computed columns run
An action cannot write; 105 lines of mutations in videos.ts
HTTP API
add_query_route derives the signature; malformed requests are 422
Five http.route blocks. Validators sit inside the function; a failure is a 500
Reactive client
None here — you would write polling or a websocket
The reason most teams pick Convex; discarded by this REST contract
Vector search result limit
No ceiling in this implementation
vectorSearch clamps to 256
Ingest, 20 videos / 10 min
62.3s, 9.7× realtime; first video 5.1s
47.6s, 12.7× realtime; first video 4.0s
Ingest, 100 videos / 63 min
361.6s, 10.45× realtime
231.5s, 16.32× realtime
Transcript search p50 (203 videos)
17.0ms
11.4ms
Reads at 203 videos, p50
5.1ms list, 0.7ms frame fetch
2.4ms list, 0.4ms frame fetch; fastest read path of the three
Search under 8 clients, p50
119.5ms; the in-process embedding serializes
65.2ms; the stateless layer stays parallel
Agent query p50 (local model)
206ms; the model runs where the data is
808ms; three compute-service round trips per question
Hosted model, 12 agent questions (paid endpoint)
12/12; 8.4s p50, the request-rate scheduler serializes pacing
12/12; 3.9s p50; zero recorded retries
Hosted swap, lines of retry code written
7; the provider scheduler paces and retries
37; pacing, Retry-After and backoff by hand
Add a derived column (lines)
1 line, 1 file
53 lines, 2 files; reverting needs a second migration
Processing for any writer
Yes — the pipeline is the schema
No, unless you add a scheduled action
Vendor checker in CI
ruff only
@convex-dev/eslint-plugin and tsc --noEmit against generated code

All three beat realtime on this laptop

Two corpus tiers: 20 videos ingested, then 100 more for 203 total. CPU, local models, not Cloud. Bold is best on that row.

Ingest

FeaturePixeltableConvex
Wall time · 20 videos62.3s47.6s
Wall time · 100 videos361.6s231.5s
Faster than realtime · 20 videos9.7x12.68x
Faster than realtime · 100 videos10.45x16.32x
Median video · 20 videos2.9s1.9s
Median video · 100 videos3.5s2.3s

Search

FeaturePixeltableConvex
Frame search p50 · 23 videos20.9ms13.9ms
Frame search p50 · 203 videos22.2ms15.3ms
Frame search p95 · 23 videos22.4ms18.1ms
Frame search p95 · 203 videos24.2ms21.5ms
Transcript search p50 · 23 videos17.9ms11.2ms
Transcript search p50 · 203 videos17.0ms11.4ms
Transcript search p95 · 23 videos21.1ms14.6ms
Transcript search p95 · 203 videos18.5ms16.5ms

Agent

Local Qwen2.5-1.5B on the 3-video baseline; retrieval is a fixed top-4 per index, so generation dominates.

FeaturePixeltableConvex
Agent query p50206.1ms807.8ms
Agent query p95235.0ms1104.3ms

Reads

FeaturePixeltableConvex
GET /videos p50 · 23 videos6.1ms1.9ms
GET /videos p50 · 203 videos5.1ms2.4ms
Frame fetch p50 · 23 videos1.7ms0.4ms
Frame fetch p50 · 203 videos0.7ms0.4ms

Under load, 8 clients

The same ten searches with eight clients in flight; measures degradation, not speed.

FeaturePixeltableConvex
Concurrent p50 · 23 videos173.6ms63.1ms
Concurrent p50 · 203 videos119.5ms65.2ms
Concurrent p95 · 23 videos264.6ms72.9ms
Concurrent p95 · 203 videos173.7ms72.8ms

Large tier ingests 20 videos (603.7s of footage) into a 3-video table; xl ingests 100 more (3,778s) for a 203-video, 7,689-frame corpus. All three finished every video at both tiers; Pixeltable logged one retried attempt at xl. Supabase led ingest at both tiers and the gap widened rather than shrank: 1.36x at large, 1.62x at xl. Search separates by single-digit milliseconds, because most of every number is embedding the query, not the index. The agent runs on the 3-video baseline, where Pixeltable is 3-4x faster because one question costs the other two three compute-service round trips and costs it none. Under eight concurrent search clients that same property inverts: the in-process embedding serializes at 120-174ms while the other two hold ~65ms. One machine, one afternoon: read the gaps, not the milliseconds.

One hosted model, three ways to call it

The agent again, with local generation swapped for nvidia/nemotron-3-super-120b-a12b on OpenRouter’s paid endpoint, identical for all three. 12 questions, 6 workers, then every patch reverts.

FeaturePixeltableConvex
Answered of 121212
Wall time36.6s18.8s
p508.4s3.9s
p9534.1s18.8s
Retriesscheduler-internal0
Lines written for the swap737

The swap is the measurement: on Pixeltable it is a 7-line schema change and the request-rate scheduler paces and retries; on the other two it is a 36-37 line helper, because pacing, Retry-After and backoff are application code. On a healthy pool that code is never exercised: all three answer every question, both loops record zero retries, so the remaining differences are who wrote the retry code and latency, where provider response time dominates and the longest tail stays on Pixeltable, whose scheduler serializes pacing. The failure that code exists to catch is real, and now measured rather than inferred: probing the same model’s free pool returned HTTP 200 carrying an upstream error and no content on 3 of 36 calls. A hand-written loop inspects the body and retries that shape; a computed column evaluates the response it is given, so a malformed 200 lands as a null answer, since the scheduler retries raised errors, not well-formed wrong ones. An earlier free-pool run recorded exactly that: five empty cells, 7/12 answers where both loops held 12/12.

Adding a column to live data

One computed title-embedding index on a populated catalog. Lines compound; these seconds do not.

FeaturePixeltableConvex
Total · 23-video catalog0.96s9.10s
Total · 123-video catalog5.22s7.66s
Schema change · 23 videos0.96s7.01s
Schema change · 123 videos5.22s6.19s
Model-free control · 23 videos0.53s1.52s
Model-free control · 123 videos0.81s1.52s
Backfill · 23 videossame step2.09s
Backfill · 123 videossame step1.48s
Lines written153
Files touched12
  • The wall-clock winner flips with corpus: Pixeltable is cheapest at 23 videos (0.96s), Supabase at 123 (3.51s). The fused step scales with rows while the script’s fixed cost amortises; a model-free control per platform separates mechanism cost from backfill work.
  • The controls expose the slope behind that flip: Pixeltable’s schema-minus-control residual grew 0.4s to 4.4s between the two corpus sizes while Convex’s stayed flat.
  • Reverting is not symmetric. Convex needs a second migration (17 of its 53 lines) because pushing a schema that no longer declares the field is rejected while documents still carry it.
  • Convex also has a lexical searchIndex path: 1 line, 2.75s at 23 videos and 1.67s at 123, no backfill, because it indexes a field that already exists rather than computing an embedding. It answers a different query.
  • After pxt schema update, an insert against the already-registered route answers 409 until pxt service update. Reads keep working. The other two resolve the table on every request.
  • Backfill time is the part that scales, and this corpus cannot show it. Pixeltable’s backfill is work proportional to the rows that changed; a backfill script is work proportional to the table.

Ingest a video

Insert a video. Frames, audio, transcripts, embeddings, and scenes have to exist after that. On Pixeltable they are the schema. On the other two they live in the ingest path and in a second service.

Pixeltable

class Videos(TableModel, name='videos'):
video: pxt.Video
title: pxt.String
audio = extract_audio(video, format='mp3')
duration_sec = pxtf.video.get_duration(video)
scenes = video.scene_detect_content(threshold=8.0)
class Frames(TableModel, name='frames', base=Videos,
iterator=frame_iterator(Videos.video, fps=1.0)):
still = pxtf.image.resize(frame, (320, 180))
__indexes__ = [pxt.EmbeddingIndex(frame, embedding=VISUAL)]
class Chunks(TableModel, name='chunks', base=Videos,
iterator=audio_splitter(Videos.audio, duration=10.0)):
transcript = transcribe(audio_segment, model='base.en').text.astype(pxt.String)
__indexes__ = [pxt.EmbeddingIndex(transcript, embedding=SEMANTIC)]
Videos.insert([{'video': 'lecture.mp4', 'title': 'CS101'}])

Convex

const { frames } = await compute("/extract-frames", { video_url, fps: FRAME_FPS });
const { embeddings } = await compute("/embed-clip", { images_b64: frames });
const frameRows = await Promise.all(frames.map(async (b64, i) => ({
frameIdx: i,
imageStorageId: await ctx.storage.store(new Blob([decodeBase64(b64)])),
embedding: embeddings[i],
})));
await ctx.runMutation(internal.videos.insertFrames, { videoId, rows: frameRows });

Search frames

Find frames of a whiteboard. Pixeltable asks the index. The other two embed the query themselves, then join or fetch rows in a second step.

Pixeltable

sim = Frames.frame.similarity(string=query)
return (
Frames.order_by(sim, asc=False)
.limit(limit)
.select(
frame_url=Frames.still,
frame_idx=Frames.pos,
video_title=Frames.title,
similarity=sim,
)
)

Convex

const { embeddings } = await compute("/embed-clip", { texts: [query] });
const hits = await ctx.vectorSearch("frames", "by_embedding", {
vector: embeddings[0],
limit: clamp(limit), // vectorSearch is 1-256
});
return await ctx.runQuery(internal.search.framesByIds, {
ids: hits.map((h) => h._id),
scores: hits.map((h) => h._score),
});

Serve over HTTP

Expose search over HTTP. Pixeltable derives the route from the query. Supabase puts five paths in one function. Convex writes five REST routes only because this contract asked for REST.

Pixeltable

api = FastAPIRouter(name='api')
api.add_insert_route(Videos, path='/videos', inputs=[Videos.video, Videos.title], background=True)
api.add_query_route(path='/videos', query=list_videos, method='get')
api.add_query_route(path='/search/frames', query=search_frames, method='post')
api.add_query_route(path='/search/transcripts', query=search_transcripts, method='post')

Convex

http.route({
path: "/search/frames",
method: "POST",
handler: httpAction(async (ctx, req) =>
guarded(async (r) => {
const body = await readJson(r);
return json(await ctx.runAction(api.search.searchFrames, {
query: requireString(body.query, "query"),
limit: readLimit(body.limit),
}));
})(req)
),
});

When to choose which platform

Choose Pixeltable when

  • The pipeline is the product

    Media in, models and retrieval out, in one Python file. Processing belongs to the table, so a row written from a shell is processed the same way as a row written over HTTP.

  • You do not want to operate a second runtime for ffmpeg

    The Convex runtime cannot execute ffmpeg. Three compute-service endpoints have no hosted-API substitute. That extra service is most of the orchestration hops.

Choose Convex when

  • Reactivity is the point

    Build the same app with Convex’s reactive client instead of five REST endpoints and http.ts disappears along with both taxes. The client re-renders on write for free, and mutations are transactional.

  • You want the easiest local backend

    npx convex dev gives a working local backend with no account and no Docker. That is the easiest install of the three, and it is not close.

Making the right choice

  • Discount the REST column

    • videos.ts (105 lines) and http.ts (90 lines) are taxes this contract imposes, not Convex’s native shape.
    • Reproduce the numbers: https://github.com/pixeltable/pixeltable-vs-supabase-vs-convex

Frequently asked questions

One file. The whole pipeline.

Declare the tables. Apply the schema. Insert a row. Serve the same file.

pip install 'pixeltable[serve]'
See how it worksGet expert guidance