Audio

What is speech-to-text?

Speech-to-text transcribes spoken audio into text that can be searched, chunked, or passed to a model.

Updated

How it works

  • The input is an audio column, or the soundtrack of a video.
  • A model writes a transcript.
  • Search and Q&A run on that transcript.

What it is not

It is not visual video search, and it is not a captions file you never index.

speech-to-text: this, and the thing it is confused with

speech-to-text: this, and the thing it is confused with
ThisNot this
InputSpeechA picture
OutputText with optional timestampsA caption of a scene
SearchThe words that were saidA face or object

Where Pixeltable fits

Pixeltable assigns a transcription computed column on pxt.Audio. The text is a normal column after that.

Questions

How does speech-to-text work?
The input is an audio column, or the soundtrack of a video. A model writes a transcript. Search and Q&A run on that transcript.
What is speech-to-text often confused with?
It is not visual video search, and it is not a captions file you never index.