Pixeltable vs Convex
Same video-intelligence app, two implementations. Pick Convex when reactivity is the point — the UI re-renders on write. Pick Pixeltable when the pipeline is the product. A REST-shaped benchmark is Convex’s worst event; discount this column accordingly.
pip install 'pixeltable[serve]'The trade
| Side | Pixeltable | Convex |
|---|---|---|
| At a glance |
|
|
What we measured
Same app. Convex wins install, transcript search, reads, concurrent load, and hosted-model latency. Pixeltable wins in-platform media, the local agent query, and the hosted swap on lines written. REST is Convex’s worst event.
| Feature | Pixeltable | Convex |
|---|---|---|
| Local install | pip install; large Python deps (torch, whisper, sentence-transformers) | npx convex dev — no account, no Docker, easiest of the three |
| Total code for this app | 129 lines, 1 file | 671 lines (425 + 246 compute-service), 7 files |
| ffmpeg, Whisper, CLIP | Computed columns, same process | compute-service; the Convex runtime cannot run them |
| Writes from an action | Insert is a row; computed columns run | An action cannot write; 105 lines of mutations in videos.ts |
| HTTP API | add_query_route derives the signature; malformed requests are 422 | Five http.route blocks. Validators sit inside the function; a failure is a 500 |
| Reactive client | None here — you would write polling or a websocket | The reason most teams pick Convex; discarded by this REST contract |
| Vector search result limit | No ceiling in this implementation | vectorSearch clamps to 256 |
| Ingest, 20 videos / 10 min | 62.3s, 9.7× realtime; first video 5.1s | 47.6s, 12.7× realtime; first video 4.0s |
| Ingest, 100 videos / 63 min | 361.6s, 10.45× realtime | 231.5s, 16.32× realtime |
| Transcript search p50 (203 videos) | 17.0ms | 11.4ms |
| Reads at 203 videos, p50 | 5.1ms list, 0.7ms frame fetch | 2.4ms list, 0.4ms frame fetch; fastest read path of the three |
| Search under 8 clients, p50 | 119.5ms; the in-process embedding serializes | 65.2ms; the stateless layer stays parallel |
| Agent query p50 (local model) | 206ms; the model runs where the data is | 808ms; three compute-service round trips per question |
| Hosted model, 12 agent questions (paid endpoint) | 12/12; 8.4s p50, the request-rate scheduler serializes pacing | 12/12; 3.9s p50; zero recorded retries |
| Hosted swap, lines of retry code written | 7; the provider scheduler paces and retries | 37; pacing, Retry-After and backoff by hand |
| Add a derived column (lines) | 1 line, 1 file | 53 lines, 2 files; reverting needs a second migration |
| Processing for any writer | Yes — the pipeline is the schema | No, unless you add a scheduled action |
| Vendor checker in CI | ruff only | @convex-dev/eslint-plugin and tsc --noEmit against generated code |
All three beat realtime on this laptop
Two corpus tiers: 20 videos ingested, then 100 more for 203 total. CPU, local models, not Cloud. Bold is best on that row.
Ingest
| Feature | Pixeltable | Convex |
|---|---|---|
| Wall time · 20 videos | 62.3s | 47.6s |
| Wall time · 100 videos | 361.6s | 231.5s |
| Faster than realtime · 20 videos | 9.7x | 12.68x |
| Faster than realtime · 100 videos | 10.45x | 16.32x |
| Median video · 20 videos | 2.9s | 1.9s |
| Median video · 100 videos | 3.5s | 2.3s |
Search
| Feature | Pixeltable | Convex |
|---|---|---|
| Frame search p50 · 23 videos | 20.9ms | 13.9ms |
| Frame search p50 · 203 videos | 22.2ms | 15.3ms |
| Frame search p95 · 23 videos | 22.4ms | 18.1ms |
| Frame search p95 · 203 videos | 24.2ms | 21.5ms |
| Transcript search p50 · 23 videos | 17.9ms | 11.2ms |
| Transcript search p50 · 203 videos | 17.0ms | 11.4ms |
| Transcript search p95 · 23 videos | 21.1ms | 14.6ms |
| Transcript search p95 · 203 videos | 18.5ms | 16.5ms |
Agent
Local Qwen2.5-1.5B on the 3-video baseline; retrieval is a fixed top-4 per index, so generation dominates.
| Feature | Pixeltable | Convex |
|---|---|---|
| Agent query p50 | 206.1ms | 807.8ms |
| Agent query p95 | 235.0ms | 1104.3ms |
Reads
| Feature | Pixeltable | Convex |
|---|---|---|
| GET /videos p50 · 23 videos | 6.1ms | 1.9ms |
| GET /videos p50 · 203 videos | 5.1ms | 2.4ms |
| Frame fetch p50 · 23 videos | 1.7ms | 0.4ms |
| Frame fetch p50 · 203 videos | 0.7ms | 0.4ms |
Under load, 8 clients
The same ten searches with eight clients in flight; measures degradation, not speed.
| Feature | Pixeltable | Convex |
|---|---|---|
| Concurrent p50 · 23 videos | 173.6ms | 63.1ms |
| Concurrent p50 · 203 videos | 119.5ms | 65.2ms |
| Concurrent p95 · 23 videos | 264.6ms | 72.9ms |
| Concurrent p95 · 203 videos | 173.7ms | 72.8ms |
Large tier ingests 20 videos (603.7s of footage) into a 3-video table; xl ingests 100 more (3,778s) for a 203-video, 7,689-frame corpus. All three finished every video at both tiers; Pixeltable logged one retried attempt at xl. Supabase led ingest at both tiers and the gap widened rather than shrank: 1.36x at large, 1.62x at xl. Search separates by single-digit milliseconds, because most of every number is embedding the query, not the index. The agent runs on the 3-video baseline, where Pixeltable is 3-4x faster because one question costs the other two three compute-service round trips and costs it none. Under eight concurrent search clients that same property inverts: the in-process embedding serializes at 120-174ms while the other two hold ~65ms. One machine, one afternoon: read the gaps, not the milliseconds.
One hosted model, three ways to call it
The agent again, with local generation swapped for nvidia/nemotron-3-super-120b-a12b on OpenRouter’s paid endpoint, identical for all three. 12 questions, 6 workers, then every patch reverts.
| Feature | Pixeltable | Convex |
|---|---|---|
| Answered of 12 | 12 | 12 |
| Wall time | 36.6s | 18.8s |
| p50 | 8.4s | 3.9s |
| p95 | 34.1s | 18.8s |
| Retries | scheduler-internal | 0 |
| Lines written for the swap | 7 | 37 |
The swap is the measurement: on Pixeltable it is a 7-line schema change and the request-rate scheduler paces and retries; on the other two it is a 36-37 line helper, because pacing, Retry-After and backoff are application code. On a healthy pool that code is never exercised: all three answer every question, both loops record zero retries, so the remaining differences are who wrote the retry code and latency, where provider response time dominates and the longest tail stays on Pixeltable, whose scheduler serializes pacing. The failure that code exists to catch is real, and now measured rather than inferred: probing the same model’s free pool returned HTTP 200 carrying an upstream error and no content on 3 of 36 calls. A hand-written loop inspects the body and retries that shape; a computed column evaluates the response it is given, so a malformed 200 lands as a null answer, since the scheduler retries raised errors, not well-formed wrong ones. An earlier free-pool run recorded exactly that: five empty cells, 7/12 answers where both loops held 12/12.
Adding a column to live data
One computed title-embedding index on a populated catalog. Lines compound; these seconds do not.
| Feature | Pixeltable | Convex |
|---|---|---|
| Total · 23-video catalog | 0.96s | 9.10s |
| Total · 123-video catalog | 5.22s | 7.66s |
| Schema change · 23 videos | 0.96s | 7.01s |
| Schema change · 123 videos | 5.22s | 6.19s |
| Model-free control · 23 videos | 0.53s | 1.52s |
| Model-free control · 123 videos | 0.81s | 1.52s |
| Backfill · 23 videos | same step | 2.09s |
| Backfill · 123 videos | same step | 1.48s |
| Lines written | 1 | 53 |
| Files touched | 1 | 2 |
- The wall-clock winner flips with corpus: Pixeltable is cheapest at 23 videos (0.96s), Supabase at 123 (3.51s). The fused step scales with rows while the script’s fixed cost amortises; a model-free control per platform separates mechanism cost from backfill work.
- The controls expose the slope behind that flip: Pixeltable’s schema-minus-control residual grew 0.4s to 4.4s between the two corpus sizes while Convex’s stayed flat.
- Reverting is not symmetric. Convex needs a second migration (17 of its 53 lines) because pushing a schema that no longer declares the field is rejected while documents still carry it.
- Convex also has a lexical searchIndex path: 1 line, 2.75s at 23 videos and 1.67s at 123, no backfill, because it indexes a field that already exists rather than computing an embedding. It answers a different query.
- After pxt schema update, an insert against the already-registered route answers 409 until pxt service update. Reads keep working. The other two resolve the table on every request.
- Backfill time is the part that scales, and this corpus cannot show it. Pixeltable’s backfill is work proportional to the rows that changed; a backfill script is work proportional to the table.
Ingest a video
Insert a video. Frames, audio, transcripts, embeddings, and scenes have to exist after that. On Pixeltable they are the schema. On the other two they live in the ingest path and in a second service.
Pixeltable
class Videos(TableModel, name='videos'):video: pxt.Videotitle: pxt.Stringaudio = extract_audio(video, format='mp3')duration_sec = pxtf.video.get_duration(video)scenes = video.scene_detect_content(threshold=8.0)class Frames(TableModel, name='frames', base=Videos,iterator=frame_iterator(Videos.video, fps=1.0)):still = pxtf.image.resize(frame, (320, 180))__indexes__ = [pxt.EmbeddingIndex(frame, embedding=VISUAL)]class Chunks(TableModel, name='chunks', base=Videos,iterator=audio_splitter(Videos.audio, duration=10.0)):transcript = transcribe(audio_segment, model='base.en').text.astype(pxt.String)__indexes__ = [pxt.EmbeddingIndex(transcript, embedding=SEMANTIC)]Videos.insert([{'video': 'lecture.mp4', 'title': 'CS101'}])
Convex
const { frames } = await compute("/extract-frames", { video_url, fps: FRAME_FPS });const { embeddings } = await compute("/embed-clip", { images_b64: frames });const frameRows = await Promise.all(frames.map(async (b64, i) => ({frameIdx: i,imageStorageId: await ctx.storage.store(new Blob([decodeBase64(b64)])),embedding: embeddings[i],})));await ctx.runMutation(internal.videos.insertFrames, { videoId, rows: frameRows });
Search frames
Find frames of a whiteboard. Pixeltable asks the index. The other two embed the query themselves, then join or fetch rows in a second step.
Pixeltable
sim = Frames.frame.similarity(string=query)return (Frames.order_by(sim, asc=False).limit(limit).select(frame_url=Frames.still,frame_idx=Frames.pos,video_title=Frames.title,similarity=sim,))
Convex
const { embeddings } = await compute("/embed-clip", { texts: [query] });const hits = await ctx.vectorSearch("frames", "by_embedding", {vector: embeddings[0],limit: clamp(limit), // vectorSearch is 1-256});return await ctx.runQuery(internal.search.framesByIds, {ids: hits.map((h) => h._id),scores: hits.map((h) => h._score),});
Serve over HTTP
Expose search over HTTP. Pixeltable derives the route from the query. Supabase puts five paths in one function. Convex writes five REST routes only because this contract asked for REST.
Pixeltable
api = FastAPIRouter(name='api')api.add_insert_route(Videos, path='/videos', inputs=[Videos.video, Videos.title], background=True)api.add_query_route(path='/videos', query=list_videos, method='get')api.add_query_route(path='/search/frames', query=search_frames, method='post')api.add_query_route(path='/search/transcripts', query=search_transcripts, method='post')
Convex
http.route({path: "/search/frames",method: "POST",handler: httpAction(async (ctx, req) =>guarded(async (r) => {const body = await readJson(r);return json(await ctx.runAction(api.search.searchFrames, {query: requireString(body.query, "query"),limit: readLimit(body.limit),}));})(req)),});
When to choose which platform
Choose Pixeltable when
- The pipeline is the product
Media in, models and retrieval out, in one Python file. Processing belongs to the table, so a row written from a shell is processed the same way as a row written over HTTP.
- You do not want to operate a second runtime for ffmpeg
The Convex runtime cannot execute ffmpeg. Three compute-service endpoints have no hosted-API substitute. That extra service is most of the orchestration hops.
Choose Convex when
- Reactivity is the point
Build the same app with Convex’s reactive client instead of five REST endpoints and http.ts disappears along with both taxes. The client re-renders on write for free, and mutations are transactional.
- You want the easiest local backend
npx convex dev gives a working local backend with no account and no Docker. That is the easiest install of the three, and it is not close.
Making the right choice
Discount the REST column
- videos.ts (105 lines) and http.ts (90 lines) are taxes this contract imposes, not Convex’s native shape.
- Reproduce the numbers: https://github.com/pixeltable/pixeltable-vs-supabase-vs-convex
Frequently asked questions
One file. The whole pipeline.
Declare the tables. Apply the schema. Insert a row. Serve the same file.