Audio

What is text-to-speech?

Text-to-speech synthesizes spoken audio from text.

Updated

How it works

  • A column holds the text.
  • A speech model writes audio.
  • The audio can be stored on the same row as the script.

What it is not

It is not transcription, and it is not cloning a private recording without a speech model.

text-to-speech: this, and the thing it is confused with

text-to-speech: this, and the thing it is confused with
ThisNot this
DirectionText to audioAudio to text
OutputA waveform or fileA transcript
Stored withThe script rowA separate media bin

Where Pixeltable fits

Pixeltable can assign a speech computed column that writes audio from a text column.

Questions

How does text-to-speech work?
A column holds the text. A speech model writes audio. The audio can be stored on the same row as the script.
What is text-to-speech often confused with?
It is not transcription, and it is not cloning a private recording without a speech model.