---
title: "What is speech-to-text?"
description: "Speech-to-text transcribes spoken audio into text that can be searched, chunked, or passed to a model."
url: "https://pixeltable.com/learn/what-is-speech-to-text"
updated: "2026-09-29"
vertical: "Audio"
doc: "https://docs.pixeltable.com/sdk/latest/openai#transcriptions"
---

# What is speech-to-text?

Speech-to-text transcribes spoken audio into text that can be searched, chunked, or passed to a model.

Updated: 2026-09-29


## On this page

- [How it works](https://pixeltable.com/learn/what-is-speech-to-text#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-speech-to-text#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-speech-to-text#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-speech-to-text#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-speech-to-text#questions)

## How it works {#how-it-works}


- The input is an audio column, or the soundtrack of a video.
- A model writes a transcript.
- Search and Q&A run on that transcript.

## What it is not {#what-it-is-not}

It is not visual video search, and it is not a captions file you never index.

## speech-to-text: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Input | Speech | A picture |
| Output | Text with optional timestamps | A caption of a scene |
| Search | The words that were said | A face or object |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable assigns a transcription computed column on pxt.Audio. The text is a normal column after that.

## Questions {#questions}

### How does speech-to-text work? {#faq-1}

The input is an audio column, or the soundtrack of a video. A model writes a transcript. Search and Q&A run on that transcript.

### What is speech-to-text often confused with? {#faq-2}

It is not visual video search, and it is not a captions file you never index.

## In the blog

- [Build a Complete Video Intelligence Pipeline in 20 Minutes](https://pixeltable.com/blog/video-intelligence-pipeline-tutorial)
- [OpenAI Whisper API Integration: Automated Audio Transcription with Pixeltable](https://pixeltable.com/blog/whisper-transcription-pixeltable)

## Related

- [Documentation](https://docs.pixeltable.com/sdk/latest/openai#transcriptions)
- [OpenAI Whisper API Integration: Automated Audio Transcription with Pixeltable](https://pixeltable.com/blog/whisper-transcription-pixeltable)
- [Audio transcription pipeline](https://pixeltable.com/use-cases/audio-transcription-pipeline)
- [audio-transcriber](https://pixeltable.com/tools/audio-transcriber)
