---
title: "What is Whisper?"
description: "Whisper is a speech-recognition model family used to transcribe audio and video soundtracks into text."
url: "https://pixeltable.com/learn/what-is-whisper"
updated: "2026-09-29"
vertical: "Audio"
doc: "https://docs.pixeltable.com/sdk/latest/openai#transcriptions"
---

# What is Whisper?

Whisper is a speech-recognition model family used to transcribe audio and video soundtracks into text.

Updated: 2026-09-29
Part of [What is speech-to-text?](https://pixeltable.com/learn/what-is-speech-to-text).

## On this page

- [How it works](https://pixeltable.com/learn/what-is-whisper#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-whisper#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-whisper#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-whisper#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-whisper#questions)

## How it works {#how-it-works}


- You pass audio and a model name.
- The result is transcript text, sometimes with segments.
- A different model embeds or answers from that text.

## What it is not {#what-it-is-not}

It is not a text embedding model, and it is not speaker diarization by itself.

## Whisper: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Job | Transcription | Embedding or chat |
| Input | Audio | A sentence |
| Speakers | Not identified unless you add that | A diarization model |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable calls Whisper through a transcription function on an audio column, for example openai.transcriptions.

## Questions {#questions}

### How does Whisper work? {#faq-1}

You pass audio and a model name. The result is transcript text, sometimes with segments. A different model embeds or answers from that text.

### What is Whisper often confused with? {#faq-2}

It is not a text embedding model, and it is not speaker diarization by itself.

## In the blog

- [OpenAI Whisper API Integration: Automated Audio Transcription with Pixeltable](https://pixeltable.com/blog/whisper-transcription-pixeltable)
- [ReelForge: Podcast Chapters and a Highlight Clip in One app.py](https://pixeltable.com/blog/reelforge-podcast-chaptering-viral-clips)
- [CallSense: Sales-Call Intelligence Without a Transcription Fleet](https://pixeltable.com/blog/callsense-sales-call-intelligence)
- [SafeStream: UGC Video Moderation as One Row, Two Signals](https://pixeltable.com/blog/safestream-ugc-video-moderation)
- [Replicate: Access Thousands of ML Models for LLMs, Images, and Audio in Pixeltable](https://pixeltable.com/blog/replicate-model-marketplace-pixeltable)
- [OpenAI GPT-4o: Complete Multimodal AI Integration with Vision, Audio, and Embeddings in Pixeltable](https://pixeltable.com/blog/openai-gpt4-multimodal-integration-pixeltable)

## Related

- [Documentation](https://docs.pixeltable.com/sdk/latest/openai#transcriptions)
- [OpenAI Whisper API Integration: Automated Audio Transcription with Pixeltable](https://pixeltable.com/blog/whisper-transcription-pixeltable)
- [OpenAI GPT-4o: Complete Multimodal AI Integration with Vision, Audio, and Embeddings in Pixeltable](https://pixeltable.com/blog/openai-gpt4-multimodal-integration-pixeltable)
- [Audio transcription pipeline](https://pixeltable.com/use-cases/audio-transcription-pipeline)
- [audio-transcriber](https://pixeltable.com/tools/audio-transcriber)
