MOSS-Speech
MOSS-Speech is a true speech-to-speech large language model without text guidance.
Category
Voice assistants
Type
Open source
Platform
Python project
Pricing
Free / open source
License
Apache-2.0
GitHub stars
139
Last updated
2026-02-13
Need the recording as text?
Whatever you record or edit with MOSS-Speech, Whipscribe turns it into an accurate, speaker-labeled transcript — 100+ languages, SRT/VTT/DOCX export, private self-hosted Whisper. About $2 per audio hour, 30 minutes free daily.
Transcribe a file →