Text-to-speech
294 tools and products, each with a live profile and alternatives list.
ElevenLabsState-of-the-art AI voices: cloning, dubbing and long-form narration with an API developers build on.Murf AIStudio-style AI voiceover tool with 120+ voices, pitch/speed control and sync-to-video for explainer content.PlayHTAI voice generation with a large voice library, voice cloning and a real-time streaming API.WellSaidEnterprise-leaning AI voiceover with consistent branded voices and team workflows for e-learning and product content.SpeechifyRead-aloud app that turns articles, PDFs and books into natural speech; popular for accessibility and study.NaturalReaderLong-standing text-to-speech tool for documents and web pages with commercial-license voices on paid tiers.MoneyPrinterTurbo利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.unslothLocal UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.GPT-SoVITS1 min voice data can also be used to train a good TTS model! (few shot voice cloning)Real-Time-Voice-CloningClone a voice in 5 seconds to generate arbitrary speech in real-timeOpenMontageWorld's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI cTTS🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and productionChatTTSA generative speech model for daily dialogue.OpenVoiceInstant voice cloning by MIT and MyShell. Audio foundation model.MockingBird🚀Clone a voice in 5 seconds to generate arbitrary speech in real-timeVoxCPMVoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloningindex-ttsAn Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech SystemCosyVoiceMulti-lingual large voice generation model, providing inference, training and deployment full-stack ability.ebook2audiobookGenerate audiobooks from e-books, voice cloning & 1158+ languages!diaA TTS model capable of generating ultra-realistic dialogue in one pass.VideoLingoNetflix-level subtitle cutting, translation, alignment, and even dubbing - one-click fully automated AI video subtitle team | Netflix级字幕切割、翻译、对齐、甚至加上配音,一键全自动视频搬supertonicLightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.edge-ttsUse Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API keyAmphionAmphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and enginTTS:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)VoiceStudioOmniVoice Studio is the Open-Source Elevenlabs alternative. AI Voice Clone, Dub, Dictate, Transcribe, Audiobook creator and Voice workflow studio.so-vits-svc-forkso-vits-svc fork with realtime support, improved interface and more features.EmotiVoiceEmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS EnginevitsVITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechMeloTTSHigh-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.espeak-ngeSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.YuEYuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but openStyleTTS2StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language ModelsAwesome-Prompt-EngineeringThis repository contains a hand-curated resources for Prompt Engineering with a focus on Generative Pre-trained Transformer (GPT), ChatGPT, PaLM etcabogenGenerate audiobooks from EPUBs, PDFs and text with synchronized captions.Kokoro-FastAPIDockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support; voice-mixing, auto-stitching, captioned timDiffSingerDiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official codeWhisperSpeechAn Open Source text-to-speech system built by inverting Whisper.metavoice-srcFoundational model for human-like, expressive TTSOpenUtauOpen singing synthesis platform / Open source UTAU successorRealtimeTTSConverts text to speech in realtimeTensorFlowTTS:stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, ChineseMOSS-TTSMOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveApplioA simple, high-quality voice conversion tool focused on ease of use and performance.TTS-WebUIA single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, elevenlabs-pythonThe official Python SDK for the ElevenLabs API.tacotronA TensorFlow implementation of Google's Tacotron speech synthesis with pre-trained model (unofficial)vall-eAn unofficial PyTorch implementation of the audio LM VALL-Eaeneasaeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)MARS5-TTSMARS5 speech model (TTS) from CAMB.AIgTTSPython library and CLI tool to interface with Google Translate's text-to-speech APIChatTTS_colab🚀 一键部署(含离线整合包)!基于 ChatTTS ,支持流式输出、音色抽卡、长音频生成和分角色朗读。简单易用,无需复杂安装。maryttsMARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure javapyttsx3Offline Text To Speech synthesis for pythonhifi-ganHiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesischat-with-gptAn open-source ChatGPT app with a voiceVieNeu-TTSVietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality • Chuyển văn bản thành giọng nói tiếng Việt • Text to speeTacotron-2DeepMind's Tacotron-2 Tensorflow implementationvall-ePyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.htmlIMS-ToucanControllable and fast Text-to-Speech for over 7000 languages!openai-edge-ttsFree, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabsdeepvoice3_pytorchPyTorch implementation of convolutional neural networks-based text-to-speech synthesis modelsclaude-code-video-toolkitAI-native video production toolkit for Claude CodeRHVoicea free and open source speech synthesizer for Russian and other languagesAuto-Synced-Translated-DubsAutomatically translates the text of a video based on a subtitle file, and then uses AI voice services to create a new dubbed & translated audio track where theread-aloudAn awesome browser extension that reads aloud webpage content with one clickGenie-TTSGPT-SoVITS ONNX Inference Engine & Model ConverterParallelWaveGANUnofficial Parallel WaveGAN (+ MelGAN & Multi-band MelGAN & HiFi-GAN & StyleMelGAN) with PytorchMiniMax-MCPOfficial MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, image generation and video generation APIs.video-podcast-makerTopic → 4K narrated video for coding agents. v4.0: all TTS via the ttsCN engine component (11 platforms incl. MiniMax voice clone, native word-level subtitle syVibeVoice-ComfyUIA comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly withSAMSoftware Automatic Mouth - Tiny Speech SynthesizerTalkingHeadTalking Head (3D): A JavaScript class for real-time lip-sync using full-body 3D avatars.Voice-Cloning-AppA Python/Pytorch app for easily synthesising human voicesOuteTTSInterface for OuteTTS models.sopranoSoprano: Instant, Ultra-Realistic Text-to-SpeechChatterbox-TTS-ServerSelf-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voiMatcha-TTS[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matchingnaturalspeech2-pytorchImplementation of Natural Speech 2, Zero-shot Speech and Singing Synthesizer, in PytorchWorldA high-quality speech analysis, manipulation and synthesis systemWavTokenizer[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modelingTwocastAI Podcast Generator for bilingual episodes, Multi Languages, Alternative to NotebookLLM;真人对话AI播客生成器,多语言,多音色audio-webuiA webui for different audio related Neural NetworksBigVGANOfficial PyTorch implementation of BigVGAN (ICLR 2023)Irodori-TTSA Flow Matching-based Text-to-Speech Model with Emoji-driven Style ControlXZVoiceFree and open source text-to-speech softwaredia2TTS model capable of streaming conversational audio in realtime.TTS-Audio-SuiteA ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, SautovcAutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder LossYourTTSYourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyonemelgan-neuripsGAN-based Mel-Spectrogram Inversion Network for Text-to-Speech SynthesisCognitive-Speech-TTSMicrosoft Text-to-Speech API sample code in several languages, part of Cognitive Services.NATSpeechA Non-Autoregressive Text-to-Speech (NAR-TTS) framework, including official PyTorch implementation of PortaSpeech (NeurIPS 2021) and DiffSpeech (AAAI 2022)NISQANISQA - Non-Intrusive Speech Quality and TTS Naturalness AssessmentStep-Audio-EditXA powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zvonage-php-sdk-coreVonage REST API client for PHP. API support for SMS, Voice, Text-to-Speech, Numbers, Verify (2FA) and more.FireRedTTSAn Open-Sourced LLM-empowered Foundation TTS SystemNaturalVoiceSAPIAdapterMake Azure natural TTS voices accessible to any SAPI 5-compatible application.flowtronFlowtron is an auto-regressive flow-based generative network for text to speech synthesis with control over speech variation and style transferEDDiscoveryCaptains log and 3d star map for Elite DangerousdiffwaveDiffWave is a fast, high-quality neural vocoder and waveform synthesizer.FastSpeechThe Implementation of FastSpeech based on pytorch.bark.cppSuno AI's Bark model in C/C++ for fast text-to-speech generationalexandria-audiobookAI-powered multi-voice audiobook generator — LLM script annotation, voice cloning, voice design, LoRA training, per-line style control, and export to MP3, chaptlueTerminal eBook Reader with Audiobook-Quality Text-to-Speech — Supports EPUB, PDF, DOCX, HTML, RTF, TXT, and MD.CloneTTSA lightweight, offline Android Text-to-Speech (TTS) engine enabling seamless system-wide voice cloning and high-fidelity text reading. / 运行在安卓本地的轻量级文字转语音 (TTS) flutter_ttsFlutter Text to Speech packageConfucius4-TTSConfucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS EnginecboardAugmentative and Alternative Communication (AAC) system with text-to-speech for the browserThorsten-VoiceThorsten-Voice: A free to use, offline working, high quality german TTS voice should be available for every project without any license struggling.MimikaStudioMimikaStudio - A local-first application for macOS (Apple Silicon) + Agentic MCP Supportbark-voice-cloning-HuBERT-quantizerThe code for the bark-voicecloning model. Training and inference.kokoro-web🔊 Kokoro Web: Free AI text-to-speech, online or self-hosted, OpenAI compatible!voicebox-pytorchImplementation of Voicebox, new SOTA Text-to-speech network from MetaAI, in Pytorchpodcast-makerFully automated video maker using motion graphics and text-to-speech synthesis to turn newsletters into daily YouTube videos.Transformer-TTSA Pytorch Implementation of "Neural Speech Synthesis with Transformer Network"chatterbox-tts-apiLocal, OpenAI-compatible text-to-speech (TTS) API using Chatterbox, enabling users to generate voice cloned speech anywhere the OpenAI API is used (e.g. Open Wedictionariez📚 A customizable dictionary extension that supports double-click lookups in 20+ languages, 1000+ dictionaries, text-to-speech, translation and Anki integration.LLaSA_trainingLLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesisvits2VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and Architecture Designf5-tts-mlxImplementation of F5-TTS in MLXxVA-SynthMachine learning based speech synthesis Electron app, with voices from specific characters from video gamestiktok-voiceSimple Python script to interact with the TikTok TTS APIPandratorTurn PDFs and EPUBs into audiobooks; subtitles or videos into dubbed videos (including translation), and more. For free. Pandrator uses local models, including mlx-serveNative LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool callingComfyUI-VibeVoiceComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audioChatterbox-TTS-ExtendedModified version of Chatterbox that accepts text files as input and no character restrictions. I use it to make audiobooks, especially for my kids.CycleGAN-VC2Voice Conversion by CycleGAN (语音克隆/语音转换): CycleGAN-VC2overlay-translator无需 ROOT 的开源 Android 屏幕实时翻译工具,适合游戏、视觉小说和漫画。支持端侧与云端 OCR、离线 LLM、多种翻译服务和文字朗读(TTS),译文可直接显示在画面上。Open-source no-root Android real-time screen translator for games, visAwesome-LLMs-meet-Multimodal-Generation🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).FlashLabs-ChromaWorlds first open-source real-time end-to-end spoken dialogue model with personalized voice cloning.orpheus-tts-localRun Orpheus 3B Locally With LM Studiovits2_pytorchunofficial vits2-TTS implementation in pytorchqwen3-tts-apple-siliconRun Qwen3-TTS text-to-speech locally on Mac (M1/M2/M3/M4). Voice cloning, voice design, custom voices. 100% offline using MLX.SummerTTSSummerTTS 是一个基于C++的独立编译的中文和英文语音合成项目,可以本地运行不需要网络,而且没有额外的依赖,一键编译完成即可用于中文和英文的语音合成。SummerTTS is a standalone Chinese and English speech synthesis(TTS) project thatpayload-aiAI Plugin is a powerful extension for the Payload CMS, integrating advanced AI capabilities to enhance content creation and management.storytellerMultimodal AI Story Teller, built with Stable Diffusion, GPT, and neural text-to-speechComfyUI-OmniVoice-TTSOmniVoice TTS nodes for ComfyUI - Zero-shot multilingual text-to-speech with voice cloning, voice design, and multi-speaker dialogueKAN-TTSKAN-TTS is a speech-synthesis training framework, please try the demos we have posted at https://modelscope.cn/models?page=1&tasks=text-to-speechknn-vcVoice Conversion With Just Nearest NeighborsStarGANv2-VCStarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice ConversionEDDICompanion application for Elite Dangerouse2-tts-pytorchImplementation of E2-TTS, "Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS", in Pytorchvixtts-demoA Vietnamese Voice Cloning Text-to-Speech Model ✨ComfyUI-VoxCPMComfyUI node for highly expressive speech and realistic zero-shot voice cloningaspeakA simple text-to-speech client for Azure TTS API.ai-video-editorOpen-source, local-first video editor where creators and AI agents edit the same real timeline.openreaderAn open-source read-along document reader server with high-quality TTS options, synchronized highlighting, and audiobook export for EPUB, PDF, DOCX, TXT, and MDAwesome-Singing-Voice-Synthesis-and-Singing-Voice-ConversionA paper and project list about the cutting edge Speech Synthesis, Text-to-Speech (TTS), Singing Voice Synthesis (SVS), Voice Conversion (VC), Singing Voice ConvVibeVoiceFusionVibeVoiceFusion is a full-stack, multi-speaker voice generation web system featuring LoRA fine-tuning, batch generation, and VRAM optimization. Based on MicrosoStyleTTSOfficial Implementation of StyleTTSAivisSpeechAivisSpeech: AI Voice Imitation System - Text to Speech SoftwareChatTTS-OpenVoiceFuse ChatTTS with OpenVoice, upload a 10-second audio clip, and clone your personalized ChatTTS voice.dl-for-emo-tts:computer: :robot: A summary on our attempts at using Deep Learning approaches for Emotional Text to Speech :speaker:soft-vcSoft speech units for voice conversiondectalkModern builds for the 90s/00s DECtalk text-to-speech application.Ming-UniAudioMing-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified RepresentationpysptkA python wrapper for Speech Signal Processing Toolkit (SPTK).talk-to-fenggeTalk to 峰哥 — 克隆任何人的声音和性格,实时语音对话,工程延迟 < 1 秒 | Clone anyone's voice & personality for real-time conversation. < 1s engineering latency.elevenlabs-jsThe official JavaScript (Node) library for the ElevenLabs API.ProDiffPyTorch Implementation of ProDiff (ACM-MM'22) with a Extremely-Fast diffusion speech synthesis pipelinealan-sdk-pcfThe Self-Coding System for Your App — Alan AI SDK for Power AppsFastDiffPyTorch Implementation of FastDiff (IJCAI'22)AssistentePessoalAssistente pessoal virtual desenvolvida com Python 🤖personalized-podcastTurn any content into a personalized AI podcast. NotebookLM-style, except you control the script, voices, and hosts. Listen in Apple Podcasts, Spotify, or any pai-skills24 cross-platform agent skills for Claude Code, Cursor, Codex & Gemini CLI — databases, messaging, research, TTS, DevOps, and Google WorkspacennmnkwiiLibrary to build speech synthesis systems designed for easy and fast prototyping.vonage-node-sdkVonage API client for Node.js. API support for SMS, Voice, Text-to-Speech, Numbers, Verify (2FA) and more.vynaroVynaro (叙影 AI) - 下一代 7 步全自动 AI 影视解说与第一人称视频创作工具 (Tauri 2 + Rust + React 19)Freeze-Omni✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLMVoiceFlow-TTS[ICASSP 2024] This is the official code for "VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching"VoxNovelVoxNovel: generate audiobooks giving each character a different voice actor.StreamingKokoroJSUnlimited text-to-speech in the Browser using Kokoro-JS, 100% local, 100% open sourceUTMOSv2UTokyo-SaruLab MOS Prediction SystemCross-Lingual-Voice-CloningTacotron 2 - PyTorch implementation with faster-than-realtime inference modified to enable cross lingual voice cloning.Dia-TTS-ServerSelf-host the powerful Dia TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), support for SafeTensors/BF16,VocelloVocello: a local, private voice studio for Apple Silicon. Write a script, pick or describe a voice, and generate speech on-device, faster than realtime on an 8 easevoice-trainerEaseVoice Trainer is a simple and user-friendly voice cloning and speech model trainer.ZeroSpeechVQ-VAE for Acoustic Unit Discovery and Voice Conversioncsm-voice-cloningSesame CSM 1B Voice CloningMsEdgeTTSA simple Azure Speech Service module that uses the Microsoft Edge Read Aloud API. https://www.npmjs.com/package/msedge-ttsGenerSpeechPyTorch Implementation of GenerSpeech (NeurIPS'22): a text-to-speech model towards zero-shot style transfer of OOD custom voice.Comprehensive-Transformer-TTSA Non-Autoregressive Transformer based Text-to-Speech, supporting a family of SOTA transformers with supervised and unsupervised duration modelings. This projecWeeaBlindA program to dub non-english media with modern AI speech synthesis, diarization, and voice cloning!VocGANVocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Networkultimate-rvcAn app for creating audio-based content such as song covers and speech using Retrieval-based Voice Conversion.T5Gemma-TTSMultilingual TTS model with voice cloning and duration control, based on T5Gemma encoder-decoder LLMmanim-voiceoverManim plugin for all things voiceoverwavegradA fast, high-quality neural vocoder.NaomiThe Naomi Project is an open source, technology agnostic platform for developing always-on, voice-controlled applications!ComfyUI-Qwen3-TTSA ComfyUI custom node suite for Qwen3-TTS, supporting 1.7B and 0.6B models, Custom Voice, Voice Design, Voice Cloning and Fine-Tuning.AutoPSTGlobal Rhythm Style Transfer Without Text TranscriptionsiSTFTNet-pytorchiSTFTNet : Fast and Lightweight Mel-spectrogram Vocoder Incorporating Inverse Short-time Fourier TransformChinese-FastSpeech2基于标贝数据继续训练,同时对原本的FastSpeech2模型做了改进,引入了韵律表征以及韵律预测模块,使中文发音更生动且富有节奏MITSUHAWorld's First Multilingual Inexpensive Therapeutic Sophisticated Ultra-responsive Holographic Agent. In simple terms, an AI you can talk to and it'll talk back IntelliNodeAccess the latest AI models like ChatGPT, LLaMA, Deepseek, Diffusion, Hugging face, and beyond through a unified prompt layer and performance evaluationKitten-TTS-ServerSelf-host the ultra-lightweight Kitten TTS model with this enhanced API server with an intuitive Web UI, large text processing for audiobooks, and GPU acceleratai_webuiAI-WEBUI: A universal web interface for AI creation, 一款好用的图像、音频、视频AI处理工具ttslearnttslearn: Library for Pythonで学ぶ音声合成 (Text-to-speech with Python)TalkieRefurbished Arduino version of the Talkie library from Peter Knight.VODERVoice Operation and Design Engine with Reproduction capabilitiesnix-tts🐤 Nix-TTS: Lightweight and End-to-end Text-to-Speech via Module-wise Distillationeasy-speech🔊 Cross browser Speech Synthesis also known as Text to speech or TTS; no dependencies; uses Web Speech APIDailyTalkOfficial repository of DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech, ICASSP 2023HiFTNetHiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transformnaturalspeech3_facodecFACodec: Speech Codec with Attribute Factorization used for NaturalSpeech 3ComfyUI-GPT_SoVITSa comfyui custom node for GPT-SoVITS! you can voice cloning and tts in comfyui nowreact-speech-kitReact hooks for Speech Recognition and Speech SynthesisvoxtreamVoXtream is a Full-Stream Zero-shot TTS model with Extremely Low Latency and Speaking rate ControlopenliveOpensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) runs locally. An openukrainian-ttsUkrainian TTS (text-to-speech) using ESPNETDDDM-VCOfficial Pytorch Implementation for "DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice CoKokoroSharpFast local TTS inference engine in C# with ONNX runtime. Multi-speaker, multi-platform and multilingual. Integrate on your .NET projects using a plug-and-play voicesmith[WIP] VoiceSmith makes training text to speech models easy.RVC-StudioThe best looking and most functional webui for RVC related tasks. See website for UI demo:Persian-tts-coquiPersian/Farsi text to speech(TTS) training using coqui ttsTensorVoxDesktop application for neural speech synthesis written in C++CoMoSpeechACM MM 2023 CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency ModelcotatronOfficial code for Cotatron @ INTERSPEECH 2020Pink-TromboneA programmable version of Neil Thapen's Pink Tromboneopenai_ttsCustom TTS component for Home Assistant. Utilizes the OpenAI speech engine or any compatible endpoint to deliver high-quality speech. Optionally offers chime anfish-audio-pythonThe official Python library for the Fish Audio API.bigvsanPytorch implementation of BigVSANml-with-audioHF's ML for Audio study groupSyntaSpeechSyntaSpeech: Syntax-aware Generative Adversarial Text-to-Speech; IJCAI 2022; Official codeMioTTS-InferenceInference server for MioTTS, a lightweight and fast LLM-based TTS model.StyleTTS-VCOfficial Implementation of StyleTTS-VCVI-SVSSinging Voice Synthesis based on VITS, different from VISingerpiper-plusMultilingual neural TTS (6 languages: JA/EN/ZH/ES/FR/PT, code supports SV) — C++, C#, Rust, Go, Python, npm (WASM). VITS + Prosody, streaming, CUDA/CoreML/DirecCross-Speaker-Emotion-TransferPyTorch Implementation of ByteDance's Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-SpeeComfyUI-VoxCPM2VoxCPM2 TTS for ComfyUI. 30 languages, voice design, controllable cloning, 48kHz audio, and LoRA trainingMeta-TTSOfficial repository of https://doi.org/10.1109/TASLP.2022.3167258. More up-to-date code is in "refactor" branch.phaseaugICASSP 2023 AcceptedlalalaiOfficial examples for LALAL.AI API: stem separation, voice cleaning, voice cloning, noise removal & audio processingkokocloneVoice Cloning, Now Inside Kokoro. Generate natural multilingual speech and clone any target voice with ease.pytorch-dc-ttsText to Speech with PyTorch (English and Mongolian)Awesome-Text-to-Speech🎤 A curated list of the latest and most influential tools, models, and resources in the Text-to-Speech sector. 🌟 Star if you like it! 🌟focalcodecA low-bitrate single-codebook 16 / 24 kHz speech codec based on focal modulationStyleTTS2🐍 🤖 Pip installable package for StyleTTS 2 human-level text-to-speech and voice cloningtext_generation_webui_xttsXTTSv2 Extension for oobabooga text-generation-webuivocal-craft-studioTuned Hindi & English AI Voice Studio 2026 – Batch Cloning & TTSvoice_clone_labClone a voice from a few minutes of audio and generate speech locally — Qwen3-TTS fine-tuning pipeline with CLI and web UIOpenToysMake Local AI Toys, Robots, Devices that work with a MacBook and an Arduino ESP32GSV-TTS-LiteGSV-TTS-Lite A high-performance inference engine specifically designed for the GPT-SoVITS text-to-speech model.(few shot voice cloning)Dia-Finetuning-VietnameseTTS Dia finetuning for VietnameseQwen3-TTS-EasyFinetuningEasy fine-tuning for Qwen3-TTS: Fast voice cloning and high-quality multilingual speech synthesis.chatterbox-finetuningFine-tuning toolkit for Chatterbox TTS & Chatterbox TURBO models. Supports 23 languages with smart vocabulary extension. Features offline preprocessing, automatvantaOpen source AI video engine built on Remotion. Voice cloning, AI avatars, animated captions, text-to-video (Wan 2.2/LTX), AI music, video editor, timeline, 100+LA-StudioLA Studio is a local-first AI audio platform for exploring, downloading, and testing speech-to-text, text-to-speech, and voice cloning modelsVideoLingo-OneClickVideoLingo+cosyvoice批量全自动视频翻译分角色配音加字幕软件,免安装一键启动整合包AceForgeAceForge is a local-first AI music workstation for Apple/OSX based on Ace-Step, DeMucs, XTTSv2viet-ttsVietTTS: An Open-Source Vietnamese Text to Speechcosyvoice-docker🎙️ CosyVoice All-in-One Docker - Production-ready TTS with Web UI, REST API & Voice Cloningpersonaplex-mlxPersonaPlex on Apple Silicon: an MLX port of NVIDIA’s full-duplex speech-to-speech model with realtime local/web modes and offline WAV inference.DEEPFAKE-AUDIO🎙️ Deepfake Audio – A neural voice cloning studio powered by SV2TTS technology.draft-to-takeDraft to Take beta: local-first AI audio production studio powered by IndexTTS2, Docker, Qwen, OmniVoice, SFX, ambience, and music sidecars.Higgs_v3-TTS-ComfyUIComfyUI nodes for higgs-audio-v3-tts-4b multilingual (100 languages) conversational TTS, zero-shot voice cloning, inline emotion/style/prosody/SFX tags, longformel-cepstral-distanceA Python library for computing the Mel-Cepstral Distance (Mel-Cepstral Distortion, MCD) between two inputs. This implementation is based on the method proposed HiggsAudio-StudioPortable Windows TTS — Higgs Audio v3 + AI text director, podcast & audiobook multi-voice. One-click, 100% offline, RU/EN, NVIDIA GPU.youtube-auto-dubLocal-first, open-source YouTube dubbing: clones the original voice into another language, time-synced. Chatterbox voice cloning + faster-whisper + NLLB, with mvibevoice-studioBeautiful voice app: record or upload to train a voice, generate speech from text or files, save & download voices.EloquentThe most feature-complete local AI workstation. Multi-GPU inference, integrated Stable Diffusion + ADetailer, voice cloning, research-grade ELO testing, and tooComfyUI-Step_Audio_EditX_TTSComfyUI nodes for Step Audio EditX - State-of-the-art zero-shot voice cloning and audio editing with emotion, style, speed control, and more.Irodori-TTS-ServerOpenAI Text-to-Speech API compatible server for Irodori-TTSGenVCSelf-supervised Generative LM-based Voice ConversionMimicManiaMimicMania is a web application that allows you to generate speech and clone voices using text-to-speech technology. With MimicMania, you can create custom voicAllVoiceLab-MCPOfficial AllVoiceLab Model Context Protocol (MCP) server, supporting interaction with powerful text-to-speech and video translation APIs.vllm-chatterbox-streamOpenAI-compatible multilingual TTS server — Chatterbox on vLLM with real-time PCM audio streaming, low time-to-first-byte (~0.7 s), voice cloning, and 23 languazerovoxzero-shot realtime TTS system, fully offline, free and open sourceNeuralTextToAudioText prompt steered synthetic audio generatorsAI-Voice-Mod-PrVoiceMod Pro is a leading real-time voice changer, soundboard, and audio processing utility designed for streamers, content creators, gamers, and vTubers. Operaspeech-studioOpen-source desktop voice-cloning studio for creators — clone a voice, script lines with emotion markers, synthesize on-device. Tauri + VoxCPM2, runs on macOS, voice-cloning-collaban improved version of Real-time-voice-cloningttslabTTSLab is THE place to easily test ANY text to text to speech model on your own pc with 0 costsupertonic3-voice-cloneTrain voice styles for Supertone/supertonic-3 model.Emoji-TTSA fork of Irodori-TTS. It includes tools for generating captions (such as with LLM-Studio) and merging models. The inference capabilities have also been enhancePocketTTS.cppSingle-file C++ TTS runtime for Pocket TTS with ONNX Runtime — voice cloning, streaming, HTTP server, FFI C APIvoice_cloneAn OpenVoice-based voice cloning tool, single executable file (~14M), supporting multiple formats without dependencies on ffmpeg, Python, PyTorch, ONNX. 基于OpenVVividDubVividDub — 公开视频语音翻译产品站;支持配音、字幕、本地化与硬字幕移除VoiceCloningGenerative voice cloning model using TTS synthesis with state-of-the-art Zero-Shot Multi-Speaker functionality. An web api built with the YourTTS TTS model to cAI-Voice-Clone-with-Coqui-XTTS-v2Free voice cloning for creators using Coqui XTTS-v2 on Google Colab. Clone your voice with just a few minutes of audio. Complete guide to build your own noteboohume-react-sdkPackages for using Hume AI and Reactmkl-vc[Interspeech 2025] Official implementation of "Training-Free Voice Conversion with Factorized Optimal Transport"voice-zeroCollection of samples suitable for use with zero-shot text to speech engines.Aurora-Audio-StudioAurora Audio Studio 1.3.0:本地优先的 Windows AI 音频工作台,提供音乐生成、配音与声音克隆、歌声转换、AI 分轨、MIDI 扒谱和视频字幕。comprehensive-bangla-ttsAiming to achieve ultimate Multilingual TTS pipeline with main focus on releasing COQUI🐸TTS(Text-to-Speech) based high performing neural voice cloning systems flangswapSelf-hosted AI video dubbing with ASR, translation, voice cloning, subtitles, and local GPU inference.PolGen-RVCПреобразование голоса на основе VITS. Ориентировано на простоту, качество и производительность.Media-AIUltimate AI Media Generation Tools Master ListbabylonBabylon.cpp is a C and C++ library for grapheme to phoneme conversion and text to speech synthesis. For phonemization a ONNX runtime port of the DeepPhonemizer awesome-voice-agentsA curated list of voice AI agent frameworks, tools, resources, and best practicesF5-TTS-Emotional-CFGZero-shot voice cloning text-to-speech (TTS) with explicit emotion class conditioning built on F5-TTSdigitaltwinUsing a single image and just 10 seconds of sample audio, our project enables you to create a video where it appears as if you're speaking the desired text.ClonedVoiceDetectionSingle- and Multi-Speaker Cloned Voice Detection: From Perceptual to Learned Features