---
title: "Audio Transcription Pipeline with OpenAI Whisper"
description: "Transcribe audio files at scale with Pixeltable and OpenAI Whisper. Automatic batching, error handling, and incremental processing."
keywords:
  - audio transcription pipeline
  - whisper api transcription
  - bulk audio transcription python
  - transcribe audio openai whisper
  - AI automation workflow
  - audio AI pipeline
complexity: "beginner"
estimated_time: "15 min"
url: "https://pixeltable.com/use-cases/audio-transcription-pipeline"
---

# Audio Transcription Pipeline with OpenAI Whisper

Transcribe audio files at scale with Pixeltable and OpenAI Whisper. Automatic batching, error handling, and incremental processing.

## Prerequisites

- Basic Python programming
- OpenAI API key

## The Problem

Transcribing large volumes of audio requires complex orchestration: file handling, API rate limiting, error retries, parallel processing, and result storage across separate systems.

## The Solution

Pixeltable automates the entire transcription workflow. Add audio files to a table, and Whisper transcription runs as a computed column with built-in batching and error handling.

## Implementation

### Audio Table

Create a table for audio files with automatic transcription.

```python
import pixeltable as pxt
from pixeltable.functions import openai

# Audio processing table
audio = pxt.create_table('app.audio', {
    'audio_file': pxt.Audio,
    'title': pxt.String,
    'speaker': pxt.String,
})

# Automatic Whisper transcription
audio.add_computed_column(
    transcript=openai.transcriptions(
        audio=audio.audio_file,
        model='whisper-1'
    )
)

# Insert files: transcription runs automatically
audio.insert([
    {'audio_file': '/recordings/meeting_01.mp3',
     'title': 'Team Standup', 'speaker': 'All'},
])
```

Every audio file is transcribed automatically on insert. Results are cached: re-inserting the same file is instant.


### Search Transcripts

Build semantic search over your transcription library.

```python
from pixeltable.functions.huggingface import sentence_transformer

# Embedding index on transcripts
audio.add_embedding_index(
    'transcript',
    string_embed=sentence_transformer.using(
        model_id='sentence-transformers/all-MiniLM-L6-v2'
    )
)

# Search meetings by content
results = audio.select(
    audio.title, audio.transcript, audio.speaker
).order_by(
    audio.transcript.similarity(string='quarterly revenue discussion'),
    asc=False
).limit(5)
```

Semantic search over transcripts lets you find relevant audio by meaning, not just keywords.


## Benefits

- 90% reduction in transcription pipeline code
- Built-in rate limiting and error handling
- Scales from single files to thousands automatically
- Incremental: only new audio files are transcribed

## Use Cases

- Podcast transcription and search
- Meeting recording analysis
- Customer call center analytics
- Lecture and training content indexing

## Performance


| Metric | Value | Description |

| --- | --- | --- |

| Throughput | 100+ files/hr | With automatic parallelization |

## Requirements

- Python 3.9+
- OpenAI API key with Whisper access

## Resources

- [AI Automation Workflow: The Pipeline Is the Table](https://pixeltable.com/blog/ai-automation-workflow) - Whisper on insert is an AI automation workflow
- [Whisper Transcription Guide](https://docs.pixeltable.com) - Detailed audio processing walkthrough