Audio HubSpeech-to-text

saa-sdk

Addressee detection for voice agents: device-directed speech detection that runs before STT, so background speech, side conversations, and the agent's own TTS echo never trigger it. No wake word, model-agnostic, drop-in for LiveKit, Pipecat, ElevenLabs, Twilio, and OpenAI. The layer your VAD and tur

Category
Speech-to-text
Type
Open source
Platform
Python project
Pricing
Free / open source
License
Apache-2.0
GitHub stars
115
Last updated
2026-07-15

GitHub repo →  ·  Website →

Need the recording as text?

Whatever you record or edit with saa-sdk, Whipscribe turns it into an accurate, speaker-labeled transcript — 100+ languages, SRT/VTT/DOCX export, private self-hosted Whisper. About $2 per audio hour, 30 minutes free daily.

Transcribe a file →

Top saa-sdk alternatives

transformers🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, fwhisper.cppPort of OpenAI's Whisper model in C/C++voiceboxThe open-source AI voice studio. Clone, dictate, create.HandyA free, open source, and extensible speech-to-text application that works completely offline.meetilyPrivacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built llamafileDistribute and run LLMs with a single file.

See the full list of saa-sdk alternatives →