---
title: "What is an audio column?"
description: "An audio column stores a waveform or audio container the schema can transcribe, embed, or synthesize from."
url: "https://pixeltable.com/learn/what-is-an-audio-column"
updated: "2026-09-29"
vertical: "Multimodal data"
doc: "https://docs.pixeltable.com/sdk/latest/openai#transcriptions"
---

# What is an audio column?

An audio column stores a waveform or audio container the schema can transcribe, embed, or synthesize from.

Updated: 2026-09-29
Part of [What is a typed media column?](https://pixeltable.com/learn/what-is-a-typed-media-column).

## On this page

- [How it works](https://pixeltable.com/learn/what-is-an-audio-column#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-an-audio-column#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-an-audio-column#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-an-audio-column#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-an-audio-column#questions)

## How it works {#how-it-works}


- The row stores the clip as audio.
- A computed column can write a transcript.
- Search can run over that transcript, or over an embedding of it.

## What it is not {#what-it-is-not}

It is not a video soundtrack that was never given a type of its own.

## audio column: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Stored | Audio the schema can open | A video file you only watch |
| Typical output | Transcript text | A waveform plot |
| Search | Words or similar speech | Filename |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable types it as pxt.Audio. A transcript assignment on that column is the usual next step.

## Questions {#questions}

### How does audio column work? {#faq-1}

The row stores the clip as audio. A computed column can write a transcript. Search can run over that transcript, or over an embedding of it.

### What is audio column often confused with? {#faq-2}

It is not a video soundtrack that was never given a type of its own.

## In the blog

- [MedDossier: Clinical Intake on One Audio + PDF + Image Row](https://pixeltable.com/blog/meddossier-clinical-intake-triage)
- [CallSense: Sales-Call Intelligence Without a Transcription Fleet](https://pixeltable.com/blog/callsense-sales-call-intelligence)
- [ClaimBot: FNOL Triage on One Heterogeneous Row](https://pixeltable.com/blog/claimbot-multimodal-fnol-triage)
- [Beyond Pandas: Why Pixeltable Is the Ultimate Tool for Multimodal Data Wrangling](https://pixeltable.com/blog/pixeltable-vs-pandas-multimodal-data-wrangling)
- [OpenAI Whisper API Integration: Automated Audio Transcription with Pixeltable](https://pixeltable.com/blog/whisper-transcription-pixeltable)

## Related

- [Documentation](https://docs.pixeltable.com/sdk/latest/openai#transcriptions)
- [OpenAI Whisper API Integration: Automated Audio Transcription with Pixeltable](https://pixeltable.com/blog/whisper-transcription-pixeltable)
- [Audio transcription pipeline](https://pixeltable.com/use-cases/audio-transcription-pipeline)
- [audio-transcriber](https://pixeltable.com/tools/audio-transcriber)
- [audio-converter](https://pixeltable.com/tools/audio-converter)
