Glossary → Voice & AI audio
Voice & AI audio

Text-to-speech (TTS)

Text-to-speech synthesizes spoken audio from written text — the inverse of speech recognition — with modern neural voices approaching natural prosody.

Contemporary TTS learned expressiveness from data rather than hand-built rules, enabling audiobook narration, screen readers and voice agents. Quality is judged on naturalness and correct emphasis, not just intelligibility.

Related terms

Frequently asked

What is Text-to-speech?

Text-to-speech synthesizes spoken audio from written text — the inverse of speech recognition — with modern neural voices approaching natural prosody.

Why does Text-to-speech matter?

Contemporary TTS learned expressiveness from data rather than hand-built rules, enabling audiobook narration, screen readers and voice agents. Quality is judged on naturalness and correct emphasis, not just intelligibility.

What terms are related to Text-to-speech?

Closely related concepts: Voice cloning, Prosody, Automatic speech recognition (ASR) — each has its own entry in this glossary.