Glossary → Voice & AI audio
Voice & AI audio
Text-to-speech (TTS)
Text-to-speech synthesizes spoken audio from written text — the inverse of speech recognition — with modern neural voices approaching natural prosody.
Contemporary TTS learned expressiveness from data rather than hand-built rules, enabling audiobook narration, screen readers and voice agents. Quality is judged on naturalness and correct emphasis, not just intelligibility.
Related terms
Voice cloningVoice cloning creates a synthetic voice that imitates a specific real person, sometimes from minutes of refere…
ProsodyProsody is the melody of speech — pitch movement, rhythm, stress and pauses — carrying meaning beyond the word…
Automatic speech recognition (ASR)Automatic speech recognition (ASR) is the technology that converts spoken language in audio into machine-reada…