Audio HubText-to-speech

VibeVoiceFusion

VibeVoiceFusion is a full-stack, multi-speaker voice generation web system featuring LoRA fine-tuning, batch generation, and VRAM optimization. Based on Microsoft's VibeVoice (AR + diffusion architecture)

Category
Text-to-speech
Type
Open source
Platform
Python project
Pricing
Free / open source
GitHub stars
489
Last updated
2026-02-23

GitHub repo →

Top VibeVoiceFusion alternatives

ElevenLabsState-of-the-art AI voices: cloning, dubbing and long-form narration with an API developers build on.Murf AIStudio-style AI voiceover tool with 120+ voices, pitch/speed control and sync-to-video for explainer content.PlayHTAI voice generation with a large voice library, voice cloning and a real-time streaming API.WellSaidEnterprise-leaning AI voiceover with consistent branded voices and team workflows for e-learning and product content.SpeechifyRead-aloud app that turns articles, PDFs and books into natural speech; popular for accessibility and study.NaturalReaderLong-standing text-to-speech tool for documents and web pages with commercial-license voices on paid tiers.

See the full list of VibeVoiceFusion alternatives →