Audio
What is speech-to-text?
Speech-to-text transcribes spoken audio into text that can be searched, chunked, or passed to a model.
Updated
How it works
- The input is an audio column, or the soundtrack of a video.
- A model writes a transcript.
- Search and Q&A run on that transcript.
What it is not
It is not visual video search, and it is not a captions file you never index.
speech-to-text: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Input | Speech | A picture |
| Output | Text with optional timestamps | A caption of a scene |
| Search | The words that were said | A face or object |
Where Pixeltable fits
Pixeltable assigns a transcription computed column on pxt.Audio. The text is a normal column after that.
Questions
- How does speech-to-text work?
- The input is an audio column, or the soundtrack of a video. A model writes a transcript. Search and Q&A run on that transcript.
- What is speech-to-text often confused with?
- It is not visual video search, and it is not a captions file you never index.