Audio HubSpeech-to-text

Speech-to-text

540 tools and products, each with a live profile and alternatives list.

transformers🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference anwhisper.cppPort of OpenAI's Whisper model in C/C++voiceboxThe open-source AI voice studio. Clone, dictate, create.HandyA free, open source, and extensible speech-to-text application that works completely offline.meetilyPrivacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local llamafileDistribute and run LLMs with a single file.faster-whisperFaster Whisper transcription with CTranslate2whisperXWhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)buzzBuzz transcribes and translates audio offline on your personal computer. Powered by OpenAI's Whisper.screenpipeYC (S26) | Record your screen 24/7 and plug into your agents. Local, private, secure. Connect to OpenClaw, Hermes agent and 100+ appsFunASROpen-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP servinpyvideotransTranslate the video from one language to another and embed dubbing & subtitles.SpeechA scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognitioleon🧠 Leon is your open-source personal assistant.kaldikaldi-asr/kaldi is the official location of the Kaldi project.vosk-apiOffline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and NodeDeepLearningExamplesState-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructsherpa-onnxSpeech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connedeep-learning-drizzleDrench yourself in Deep Learning, Reinforcement Learning, Machine Learning, Computer Vision, and NLP by learning from these exciting lectures!!PaddleSpeechEasy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verificationspeech-to-speechBuild local voice agents with open-source modelsvoice-proGradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processispeechbrainA PyTorch-based Speech ToolkitopenvinoOpenVINO™ is an open source toolkit for optimizing and deploying AI inferencepyannote-audioNeural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embeddingRTranslatorOpen source real-time translation app for Android that runs locallyNoiseTorchReal-time microphone noise suppression on Linux.RealtimeSTTA robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.silero-vadSilero VAD: pre-trained enterprise-grade Voice Activity DetectorespnetEnd-to-End Speech Processing ToolkitinferenceSwap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — aSenseVoiceOpen-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.speech_recognitionSpeech recognition module for Python, supporting several engines and APIs, online and offline.ASRT_SpeechRecognitionA Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统youtube-transcript-apiThis is a python API which allows you to get the transcript/subtitles for a given YouTube video. It also works for automatically generated subtitles and it doesmlx-audioA text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silwukong-robot🤖 wukong-robot 是一个简单、灵活、优雅的中文语音对话机器人/智能音箱项目,支持ChatGPT多轮对话能力,还可能是首个支持脑机交互的开源智能音箱项目。vibeTranscribe on your own!annyang💬 Speech recognition for your sitewav2letterFacebook AI Research's Automatic Speech Recognition Toolkitargmax-oss-swiftOn-device Speech AI for Apple SiliconPaddleXAll-in-One Development Tool based on PaddlePaddleFunClipFunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.silero-modelsSilero Models: pre-trained text-to-speech models made embarrassingly simplecactusQuantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.Recorderhtml5 js 录音 mp3 wav ogg webm amr g711a g711u 格式,支持pc和Android、iOS部分Web浏览器、Hybrid App(提供Android iOS App源码)、微信,提供ASR语音识别转文字 H5版语音通话聊天示例 DTMF编码解码whisper-diarizationAutomatic Speech Recognition with Speaker Diarization based on OpenAI WhisperopenwhisprVoice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.dograhOpen source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builderYouDub-webui开源 AI 视频本地化工具:自动完成 YouTube/Bilibili 视频下载、字幕识别与翻译、语音克隆配音、音轨混合和字幕压制。wenetProduction First and Production Ready End-to-End Speech Recognition ToolkitporcupineOn-device wake word detection powered by deep learningml-roadMachine Learning and Agentic AI Resources, Practice and ResearchsttVoice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式whisper-jaxJAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.SmartSub视频转字幕、字幕翻译、AI 配音与声音克隆、字幕烧录——免费开源的一站式桌面工具。基于 Whisper / FunASR 等本地模型离线语音转文字,批量处理 + 全平台 GPU 加速,跨 Windows / macOS / Linux。Free, open-source desktop app to generate,fastrtcThe python library for real-time communicationAI-Youtube-Shorts-GeneratorOpen-source alternative to Opus Clip, Vidyo.ai, Klap & SubMagic. Turn long-form YouTube videos into viral 9:16 shorts using LLM highlight detection, Whisper trapocketsphinxA small speech recognizercheetahMac app for crushing tech interviews with AIWhisperLiveA nearly-live implementation of OpenAI's Whisper.ODSTurn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.distil-whisperDistilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.LLPlayerThe media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!auto-subsOn-device subtitle generation that connects directly to DaVinci Resolve, Premiere, and After Effects.embarkFramework for serverless Decentralized Applications using Ethereum, IPFS and other platformsStreamer-SalesStreamer-Sales 销冠 —— 卖货主播 LLM 大模型🛒🎁,一个能够根据给定的商品特点从激发用户购买意愿角度出发进行商品解说的卖货主播大模型。🚀⭐内含详细的数据生成流程❗ 📦另外还集成了 LMDeploy 加速推理🚀、RAG检索增强生成 📚、TTS文字转语音🔊、数字人生成 🦸、 Agent 使用网络查询实时LiveCaptions-TranslatorLightweight and powerful real-time audio/speech translation tool based on Windows LiveCaptions.chatgpt-telegram-bot🤖 A Telegram bot that integrates with OpenAI's official ChatGPT APIs to provide answers, written in Pythonchatgpt-javaChatGPT Java SDK支持流式输出、Gpt插件、联网。支持OpenAI官方所有接口。ChatGPT的Java客户端。OpenAI GPT-3.5-Turb GPT-4 Api Client for Javawhisper-webML-powered speech recognition directly in your browserwhisper-asr-webserviceOpenAI Whisper ASR Webservice APIruby-openaiOpenAI API + Ruby! 🤖❤️ GPT-5 & Realtime WebRTC compatible!whisper-standalone-winWhisper & Faster-Whisper standalone executables for those who don't want to bother with Python.BayLing-SpeechLLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the awesome-speech-recognition-speech-synthesis-papersAutomatic Speech Recognition (ASR), Speaker Verification, Speech Synthesis, Text-to-Speech (TTS), Language Modelling, Singing Voice Synthesis (SVS), Voice Conve3D-SpeakerA Repository for Single- and Multi-modal Speaker Verification, Speaker Recognition and Speaker DiarizationwillowOpen source, local, and self-hosted Amazon Echo/Google Home competitive Voice Assistant alternativewhishperTranscribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!openlessHold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input for macOS & Windows. (按住快捷键说话,松开即得润色后的文字)openai.NET library for the OpenAI service API by Betalgo Ranulfaster-whisper-GUIfaster_whisper GUI with PySide6Whisper-WebUIA Web UI for easy subtitle using whisper model.whisper-timestampedMultilingual Automatic Speech Recognition with word-level timestamps and confidenceAutomatic_Speech_RecognitionEnd-to-end Automatic Speech Recognition for Madarian and English in TensorflowvexaOpen-source meeting transcription API for Google Meet, Microsoft Teams & Zoom. Auto-join bots, real-time WebSocket transcripts, MCP server for AI agents. Self-hFluidAudioFrontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open soSTT🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.pluelyThe Open Source Alternative to Cluely - A lightning-fast, privacy-first AI assistant that works seamlessly during meetings, interviews, and conversations withouawesome-whisper🔊 Awesome list for Whisper — an open-source AI-powered speech recognition system developed by OpenAIGPA[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!auto-subtitleAutomatically generate and overlay subtitles for any video.ququ开源免费的 Wispr Flow 替代方案 | 集成FunASR本地模型和可配置大语言模型的下一代中文桌面语音工作流AudioNotes快速提取音视频内容,整理成一份结构化的markdown笔记ten-vadVoice Activity Detector (VAD) : low-latency, high-performance and lightweightvoice_datasets🔊 A comprehensive list of open-source datasets for voice and sound computing (95+ datasets).kalosmInstant, controllable, local pre-trained AI models in Rustyoutube-shorts-pipelineAutomated YouTube Shorts pipeline: news → script → AI visuals → voiceover → captions → uploadtensorflow-speech-recognition🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networkssoloudFree, easy, portable audio engine for gamesqwen-audio-agentA realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI AgentsWhisperJAVASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAVOnnxStreamLightweight inference library for ONNX files, written in C++. It can run Stable Diffusion XL 1.0 on a RPI Zero 2 (or in 298MB of RAM) but also Mistral 7B on desvadVoice activity detector (VAD) for the browser with a simple APIclaude-real-videoLet Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.diartA python package to build AI-powered real-time audio applicationsparlorOn-device, real-time multimodal AI with features similar to GPT-Livemasr中文语音识别; Mandarin Automatic Speech Recognition;FireRedASROpen-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also ofjuliusOpen-Source Large Vocabulary Continuous Speech Recognition Enginelip-reading-deeplearning:unlock: Lip Reading - Cross Audio-Visual Recognition using 3D Architecturesawesome-diarizationA curated list of awesome Speaker Diarization papers, libraries, datasets, and other resources.audapolisan editor for spoken-word audio with automatic transcriptionalan-sdk-iosThe Self-Coding System for Your App — Alan AI SDK for iOSopenai-kotlinOpenAI API client for Kotlin with multiplatform and coroutines capabilities.alan-sdk-androidThe Self-Coding System for Your App — Alan AI SDK for AndroidTuyaOpenNext-gen AI+IoT framework for T2/T3/T5AI/ESP32/and more – Fast IoT and AI Agent hardware integrationwhisper-turboCross-Platform, GPU Accelerated Whisper 🏎️transcribe.cppggml speech-to-text inference for 16+ model familieskalliopeKalliope is a framework that will help you to create your own personal assistant.sherpa-ncnnReal-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, alan-sdk-flutterThe Self-Coding System for Your App — Alan AI SDK for Flutterbailing百聆 是一个类似GPT-4o的语音对话机器人,通过ASR+LLM+TTS实现,集成DeepSeek R1等优秀大模型,接入openClaw,真正的个人语音助手,时延低至800ms,Mac等低配置也可运行,支持打断typewhisper-macLocal speech-to-text for macOS on-device AI, fully private, optional cloudreact-native-executorchDeclarative way to run AI models in React Native on device, powered by ExecuTorch.subsai🎞️ Subtitles generation tool (Web-UI + CLI + Python package) powered by OpenAI's Whisper and its variants 🎞️alan-sdk-ionicThe Self-Coding System for Your App — Alan AI SDK for IonicamicaAmica is an open source interface for interactive communication with 3D characters with voice synthesis and speech recognition.dsnoteSpeech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.obs-localvocalOBS plugin for local speech recognition and captioning using AIopenscreenRecord your screen, ship a demo. Free and open-source, GPU-accelerated, no watermarks, no subscriptions. Windows, macOS, Linux. Actively maintained.RCLITalk to your Mac, query your docs, no cloud required. On-device voice AI + RAGvideo-analyzerAnalyze videos using LLMs, Computer Vision and Automatic Speech RecognitionComfyUI_Custom_Nodes_AlekPetCustom nodes that extend the capabilities of Comfyuivllm-mlxHigh-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, aSALMONNSALMONN family: A suite of advanced multi-modal LLMsamical🎙️ AI Dictation App - Open Source and Local-first ⚡ Type 3x faster, no keyboard needed. 🆓 Powered by open source models, works offline, fast and accurate.ai-dev-galleryAn open-source project for Windows developers to learn how to add AI with local models and APIs to Windows apps.Fun-ASROpen-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.SpeechT5Unified-Modal Speech-Text Pre-Training for Spoken Language Processingyt-whisperUsing OpenAI's Whisper to automatically generate YouTube subtitlesminutesEvery meeting, every idea, every voice note, searchable by your AI. Open-source, privacy-first conversation memory layer.Speech-AI-Forge🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.Dragonfirethe open-source virtual assistant for Ubuntu based Linux distributionsSpeech-Emotion-AnalyzerThe neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)SoniTranslateSynchronized Translation for Videos. Video dubbingnightingaleMachine learning powered Karaoke app (with scores!)open-speech-corpora💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologiesmlx-tuneFine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.wespeakerResearch and Production Oriented Speaker Verification, Recognition and Diarization Toolkitwhisper-ctranslate2Whisper command line client compatible with original OpenAI client based on CTranslate2.gp.nvimGp.nvim (GPT prompt) Neovim AI plugin: ChatGPT sessions & Instructable text/code operations & Speech to text [OpenAI, Ollama, Anthropic, ..]voicemodeNatural voice conversations with Claude CodeJ.A.R.V.I.SPersonal Assistant built using python libraries. It does almost anything which includes sending emails, Optical Text Recognition, Dynamic News Reporting at any airunnerOffline inference engine for art, real-time voice conversations, LLM powered chatbots and automated workflowsVideoChat实时交互数字人,可自定义形象与音色,支持音色克隆,对话延迟低至3s。Real-time voice interactive digital human, customizable appearance and voice, supporting voice cloning, with initial package dCrisperWhisperVerbatim Automatic Speech Recognition with improved word-level timestamps and filler detectionStreamSpeechStreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.MToolsMTools 是一个功能强大的多功能桌面应用程序,集成了音视频处理、图片编辑、文本操作和编码工具,内置AI增强功能。旨在简化您的工作流程,提升生产效率prunaPruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead.artyom.jsA voice control - voice commands - speech recognition and speech synthesis javascript library. Create your own siri,google now or cortana with Google Chrome witwhisperWhisper is a file-based time-series database format for Graphite.vosk-serverWebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi librarieskubeaiAI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.Whisper-FinetuneFine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelclaude-video-visionGive Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysistrussThe simplest way to serve AI/ML models in productionTreasure-of-Transformers💁 Awesome Treasure of Transformers Models for Natural Language processing contains papers, videos, blogs, official repo along with colab Notebooks. 🛫☑️Foundation-Models-Framework-LabA practical lab for building, testing, and evaluating apps with Apple's Foundation Models framework.video-subtitle-generator视频音频生成字幕,生成srt文件。无需申请第三方API,本地实现音频转文本。基于Transformer的视频字幕生成框架。A GUI tool for generating subtitle from videos and generating srt files.lottiA private logbook with a staff of personal AI assistants. Agents read what you record and propose what to do next — you approve the changes. End-to-end encryptedc_ttsA TensorFlow Implementation of DC-TTS: yet another text-to-speech modelhyprwhsprNative speech-to-text for Linux - Fast, accurate, private, and hackable system-wide dictationIrene-Voice-AssistantИрина - русский голосовой ассистент для работы оффлайн. Поддерживает скиллы через плагины.lhotseTools for handling multimodal data in machine learning projects.binioua self-hosted webui for 30+ generative aialan-sdk-cordovaThe Self-Coding System for Your App — Alan AI SDK for Cordovaconformer[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)Mega-ASRFirst foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every otspeech-swiftAI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreMLAI-Waifu-VtuberAI Vtuber for Streaming on Youtube/Twitchwhisper-writer💬📝 A small dictation app using OpenAI's Whisper speech recognition model.kaldi-gstreamer-serverReal-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.WhisperboardThe open-source iOS app that's making quality voice transcription more accessible on mobile devices.voxtypeVoice-to-text with push-to-talk for Wayland compositorsvosk-android-demoOffline speech recognition for Android with Vosk library.pykaldiA Python wrapper for KaldiviolinOpen-source Video Translation SkillPatterOpen-source voice-AI SDK. The Vapi/Retell alternative for builders who want to own the stack. Give your AI agent a phone number in 4 lines — Python and TypeScriPython-ai-assistantPython AI assistant 🧠TensorFlowASR:zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwordsStoryToolkitAIAn editing tool that uses AI to transcribe, understand content and search for anything in your footage, integrated with ChatGPT and other AI modelsvoquillOpen source voice dictation technologymurmureFully local, private and cross platform Speech-to-Text with LLM Post-processingsherpaSpeech-to-text server framework with next-gen Kaldiathenaan open-source implementation of sequence-to-sequence based speech processing enginemycroft-preciseA lightweight, simple-to-use, RNN wake word listenerVoiceStreamAINear-Realtime audio transcription using self-hosted Whisper and WebSocket in Python/JSaudio-ai-hubThe hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.transcriptionstreamturnkey self-hosted offline transcription and diarization service with llm summarybotium-speech-processingBotium Speech ProcessingespressoEspresso: A Fast End-to-End Neural Speech Recognition Toolkitwhisper.netWhisper.net. Speech to text made simple using Whisper ModelsmuesliMuesli - local meeting transcription + dictation for macOS (Granola + WisprFlow alternative)Uncensored-Local-StudioUncensored local AI studio for Windows, Linux, and macOS. Zero-setup GUI for Image Generation, GGUF LLMs, Text to Speech & Speech to TextjiwerEvaluate your speech-to-text system with similarity measures such as word error rate (WER)whisper.apiThis project provides an API with user level access support to transcribe speech to text using a finetuned and processed Whisper ASR model.voicy@voicybot Telegram bot main repositoryinaSpeechSegmenterCNN-based audio segmentation toolkit. Allows to detect speech, music, noise and speaker gender. Has been designed for large scale gender equality studies based vocotype-cliVocoType 是一款运行在本地端侧的隐私安全语音输入工具,通过快捷键即可将语音实时转换为文字并自动输入到当前应用。支持语音转文字MCP、AI 优化文本、自定义替换词典、录音视频转文字等功能,让语音输入更高效、更安全。SmartJavaAI🔥🔥🔥Java免费离线AI算法工具箱,支持人脸识别,活体检测,表情识别、目标检测、实例分割、行人检测、OCR文字识别、车牌识别、表格识别、ASR+TTS、机器翻译等功能,Maven引用即可使用。支持PyTorch、Tensorflow,已集成 Mtcnn、InsightFace、SeetaFace6、YOLOv8~v1TheWhisperOptimized Whisper models for streaming and on-device usenonoCAPTCHAAn asynchronized Python library to automate solving ReCAPTCHA v2 using audioTypeNoA free, open source, privacy-first voice input app for macOS.TwitchLibC# Twitch Chat, Whisper, API and PubSub Library. Allows for chatting, whispering, stream event subscription and channel/account modification. Supports everythispeechpy:speech_balloon: SpeechPy - A Library for Speech Processing and Recognition: http://speechpy.readthedocs.io/en/latest/local-talking-llmA talking LLM that runs on your own computer without needing the internet.Jarvis-Desktop-Voice-AssistantA python based desktop voice assistant capable of executing system-level commands, integrating speech recognition and text-to-speech, and handling asynchronous PPASR基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型Easy-Voice-ToolkitA user-friendly toolkit for voice recgonition/transcription/conversion etc. | 简单易用的语音工具箱subvertGenerate subtitles, summaries, and chapters from videos in secondsVideoSubtitleGenerator批量为本地视频生成字幕文件,并可将字幕文件翻译成其它语言, 跨平台支持 window, mac 系统auditokAn voice activity detection and audio segmentation toolDeepSpeech-examplesExamples of how to use or integrate DeepSpeechGLM-ASRGLM-ASR-Nano: A robust, open-source speech recognition model with 1.5B parametersOpenAI-UnityAn unofficial OpenAI Unity Package that aims to help you use OpenAI API directly in Unity Game engine.react-speech-recognition💬Speech recognition for your React appCTCDecoderConnectionist Temporal Classification (CTC) decoding algorithms: best path, beam search, lexicon search, prefix search, and token passing. Implemented in Pythonwhisper-playgroundBuild real time speech2text web apps using OpenAI's Whisper https://openai.com/blog/whisper/go-carbonGolang implementation of Graphite/Carbon server with classic architecture: Agent -> Cache -> Persisterwhisper-flowWhisper-Flow is a framework designed to enable real-time transcription of audio content using OpenAI’s Whisper model. Rather than processing entire files after kurDescriptive Deep Learningvoxtral-mini-realtime-rsVoxtral ASR & TTS running natively and in the browser. A Rust implementation of Mistral's Voxtral mini realtime ASR / TTS using the Burn ML frameworkSpeech-TransformerA PyTorch implementation of Speech Transformer, an End-to-End ASR with Transformer network on Mandarin Chinese.generate-subtitlesGenerate transcripts for audio and video content with a user friendly UI, powered by Open AI's Whisper with automatic translations and download videos automaticwhisper.rnReact Native binding of whisper.cpp.lobe-tts🎤 Lobe TTS - A high-quality & reliable TTS/STT library for Server and Browsersglang-omniSGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.TTS-Voice-WizardSpeech to Text to Speech. Song now playing. Sends text as OSC messages to VRChat to display on avatar. (STTTS) (Speech to TTS) (VRC STT System) (VTuber TTS)MioSub一站式全自动字幕生成软件,下载、转录、翻译、压制全流程覆盖,无需人工介入 / One-stop automated subtitle generator. Handles downloading, transcription, translation, and hardcoding—zero human intervejuneLocal voice chatbot for engaging conversations, powered by Ollama, Hugging Face Transformers, and Coqui TTS Toolkitwhisper_micProject that allows one to use a microphone with OpenAI whisper.use-whisperReact hook for OpenAI Whisper with speech recorder, real-time transcription, and silence removal built-inOpenCluelyOpenCluely is a free, open source Cluely (alternative), built for technical interviews like DSA, OAs, and CP. It offers an invisible overlay, real-time AI help,SwiftWhisper🎤 The easiest way to transcribe audio in SwiftZerolanLiveRobotAI VTuber with LLM, ASR, TTS, OCR, CV and more technologies to live stream or play Minecraft with you.voxt🎙️ An intelligent voice productivity assistant that turns speech into clean text, useful actions, and structured knowledge. It helps users capture ideas, communNotelyVoiceA 100% private AI voice transcription app that converts speech to text in 100+ languages. Built with Compose Multiplatform for Android & iOS using Whisper AI - cn2an📦 快速转化「中文数字」和「阿拉伯数字」~ (最新特性:分数,日期、温度等转化)viral-clips-crewYour CrewAI Powered Video Editing AssistantPaddlePaddle-DeepSpeech基于PaddlePaddle实现的语音识别,中文语音识别。项目完善,识别效果好。支持Windows,Linux下训练和预测,支持Nvidia Jetson开发板预测。dlaDeep learning for audio processingmlx-audio-swiftA modular Swift SDK for audio processing with MLX on Apple SiliconchaplinA real-time silent speech recognition tool.whisper.unityRunning speech to text model (whisper.cpp) in Unity3d on your local machine.vocalinuxFree, open-source, 100% offline voice dictation for Linux. Speak and type anywhere via whisper.cpp, Whisper & VOSK engines, GPU-accelerated, works on X11 + WaylvuiReal-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× rGigaAMFoundational Model for Speech Recognition TasksBiBi-Keyboard说点啥(BiBi Keyboard):一个基于 Kotlin 的 Android 平台的 LLM 与 ASR 语音输入法键盘应用 An LLM ASR voice input method keyboard application for the Android platform based on KotlinallosaurusAllosaurus is a pretrained universal phone recognizer for more than 2000 languageschinese_text_normalizationChinese text normalization for speech processingawesome-large-audio-modelsCollection of resources on the applications of Large Language Models (LLMs) in Audio AI.bolnaConversational voice AI agentsTranscribroPrivate and on-device speech recognition keyboard and service for Android.MASRPytorch实现的流式与非流式的自动语音识别框架,同时兼容在线和离线识别,目前支持Conformer、Squeezeformer、DeepSpeech2模型,支持多种数据增强方法。cursesSpeech to Text and KB input captions for OBS, VRChat, Twitch chat and DiscordopenspeechOpen-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.voice-aiRapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, mulRecapOpen Source, Privacy-First, macOS-Native AI Meeting SummaryrhinoOn-device Speech-to-Intent engine powered by deep learningspeech-to-text-benchmarkspeech to text benchmark frameworkbulk_transcribe_youtube_videos_from_playlistEasily take an entire YouTube playlist and turn it into high quality transcripts using Whisper.whisper_androidOffline Speech Recognition with OpenAI Whisper and TensorFlow Lite for AndroidINTERSPEECH-2023-24-PapersINTERSPEECH 2023-2024 Papers: A complete collection of influential and exciting research papers from the INTERSPEECH 2023-24 conference. Explore the latest advafree-spoken-digit-datasetA free audio dataset of spoken digits. An audio version of MNIST.openlrcTranscribe and translate voice into LRC file using Whisper and LLMs (GPT, Claude, et,al). 使用whisper和LLM(GPT,Claude等)来转录、翻译你的音频为字幕文件。hearCommand line interface for the built-in speech recognition and transcription capabilities in macOS.cheetahOn-device streaming speech-to-text engine powered by deep learningexpo-speech-recognitionSpeech Recognition for React Native Expo projectsSpeech-TranslateA realtime speech transcription and translation application using Whisper OpenAI and free translation API. Interface made using Tkinter. Code written fully in POwlA personal wearable AI that runs locallyTranscriptionSuiteA fully local and private Speech-To-Text app, offering multiple model backends, diarization & calendar mode - Available for Windows, macOS & Linuxlora-svcsinging voice change based on whisper, and lora for singing voice clonegpt-homeChatGPT at home! A better alternative to commercial smart home assistants, built on the Raspberry Pi using LiteLLM and LangGraph.sonus:speech_balloon: /so.nus/ STT (speech to text) for Node with offline hotword detectionFireRedASR2SA SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, codekospeechOpen-Source Toolkit for End-to-End Korean Automatic Speech Recognition leveraging PyTorch and Hydra.Proctoring-AICreating a software for automatic monitoring in online proctoringwhisperIMEAndroid Input Method Editor (IME) based on Whisperspeech-to-textReal-time transcription using faster-whisperoptimum-intel🤗 Optimum Intel: Accelerate inference with Intel optimization toolsRapidASR📣 商用级开源语音自动识别程序库,开箱即用,全平台支持,中英文混合识别。A Cross-platform implementation of ASR inference. It's based on ONNXRuntime and FunASR. We provide a set of easier APIs to agentchainChain together LLMs for reasoning & orchestrate multiple large models for accomplishing complex tasksSpeech-BackbonesThis is the main repository of open-sourced speech technology by Huawei Noah's Ark Lab.pindropA native macOS menu bar dictation app using local speech-to-text with WhisperKitswiftFast voice assistant powered by Groq, Cartesia, and Vercel.echokit_serverOpen Source Voice Agent PlatformVLog[CVPR 2025] Video Narration as Vocabulary & Video as Long DocumentCTCWordBeamSearchConnectionist Temporal Classification (CTC) decoder with dictionary and language model.SayboardAn open-source on-device voice IME (keyboard) for Android using the Vosk library.WhisperS2TAn Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference EngineWhisperSubTranslateA free, local desktop app to extract subtitles (SRT) from video and translate them into any language — unlimited use, no signup, no cloud.alan-sdk-reactnativeThe Self-Coding System for Your App — Alan AI SDK for React Nativeauto-captionA cross-platform real-time subtitle display software. 一个跨平台的实时字幕显示软件。Playwright-reCAPTCHAA Python library for solving reCAPTCHA v2 and v3 with PlaywrightVRCOSCA modular node-programming language, program creator, animation system, toolkit, router, and debugger made for VRChatmacparakeetFast, private, local-first voice app for Apple Silicon Macs — dictation, file/media transcription, meeting recording, Transforms, and a public automation CLI. Fvoice-overlay-ios🗣 An overlay that gets your user’s voice permission and input as text in a customizable UIFastASR这是一个用C++实现ASR推理的项目,它依赖很少,安装也很简单,推理速度很快,在树莓派4B等ARM平台也可以流畅的运行。 支持的模型是由Google的Transformer模型中优化而来,数据集是开源wenetspeech(10000+小时)或阿里私有数据集(60000+小时), 所以识别效果也很好,可以媲美许多商用的SpectralClusterPython re-implementation of the (constrained) spectral clustering algorithms used in Google's speaker diarization papers.LeaderboardSpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.decipherEffortlessly add AI-generated transcription subtitles to your videosAwesome-Korean-Speech-Recognition한국어 음성인식 STT API 리스트. 각 성능 벤치마크.CrispASRC++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal forced alignment, and moreCleanS2SHigh-quality and streaming Speech-to-Speech interactive agent in a single file. 只用一个文件实现的流式全双工语音交互原型智能体!vosk-browserA speech recognition library running in the browser thanks to a WebAssembly build of VoskparrotsAutomatic Speech Recognition(ASR), Text-To-Speech(TTS) engine. 中英语音识别、多角色语音合成,支持多语言,准确率高ICASSP-2023-24-PapersICASSP 2023-2024 Papers: A complete collection of influential and exciting research papers from the ICASSP 2023-24 conferences. Explore the latest advancements Audar-ASR-V1Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.ollama-voice-macMac compatible Ollama Voiceai-pronunciation-trainerThis tool uses AI to evaluate your pronunciation.willow-inference-serverOpen source, local, and self-hosted highly optimized language inference server supporting ASR/STT, TTS, and LLM across WebRTC, REST, and WSLiveTranslateReal-time audio translation, captures system audio + mic, runs ASR (Whisper/SenseVoice), translates via LLM API with streaming display. Perfect for VTubers, ltranscribeeopen source audio and video transcription softwareScribeWizardScribeWizard: Generate organized notes from audio using Groq, Whisper, and Llama3android-vadAndroid Voice Activity Detection (VAD) library. Supports WebRTC VAD GMM, Silero VAD DNN, Yamnet VAD DNN models.FireRedVADA SOTA Industrial-Grade Voice Activity Detection & Audio Event Detection, supporting 100+ languages, outperforming Silero-VAD, TEN-VAD, FunASR-VAD and WebRTC-VAwhisplay-ai-chatbotPocket-sized AI chatbot built using a RPI Zero 2w / 5libfaceidlibfaceid is a research framework for prototyping of face recognition solutions. It seamlessly integrates multiple detection, recognition and liveness models w/UniSpeechUniSpeech - Large Scale Self-Supervised Learning for SpeechaudioflareAn all-in-one AI audio playground using Cloudflare AI Workers to transcribe, analyze, summarize, and translate any audio file.leopardOn-device speech-to-text engine powered by deep learningltuCode, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".react-micRecord audio from a user's microphone and display a cool visualization.Fast-Powerful-Whisper-AI-Services-API⚡ 一款用于自动语音识别 (ASR)、翻译的高性能异步 API。不需要购买Whisper API,使用本地运行的Whisper模型进行推理,并支持多GPU并发,针对分布式部署进行设计。还内置了包括TikTok、抖音等社交媒体平台的爬虫,可实现来自多个社交平台的无缝媒体处理,为媒体内容数据自动化处理提供了强大且可扩展的解huggingsoundHuggingSound: A toolkit for speech-related tasks based on Hugging Face's toolsspeech_datasetThe dataset of Speech RecognitionMLX-Auto-Subtitled-Video-GeneratorGenerate accurate transcripts using Apple's MLX frameworkBiliSum为 Bilibili、YouTube 及本地视频提供 AI 视频摘要和知识库.AI video summarizer and knowledge base for Bilibili, YouTube and local videos.TeroSubtitlerTero Subtitler is an open source, cross-platform, and free subtitle editing software.talking-avatar-with-aiThis project is a digital human that can talk and listen to you. It uses OpenAI's GPT to generate responses, OpenAI's Whisper to transcript the audio, Eleven Ladeepgram-python-sdkOfficial Python SDK for Deepgram.docker-whisperXDockerfile for WhisperX: Automatic Speech Recognition with Word-Level Timestamps and Speaker Diarization (Dockerfile, CI image build and test)autoEdit_2Fast text based video editing, node Electron Os X desktop app, with Backbone front end.speaker-idThis repository contains audio samples and supplementary materials accompanying publications by the "Speaker, Voice and Language" team at Google.video-recap-skillsClip any video into a narration recap with claude code skill|用 claude code skill 把任何视频剪辑成中文解说视频,支持剪映导出whisper-guiA simple GUI to use Whisper.opentypelessOpen-source AI voice typing for macOS, Windows, and Linux. Press a hotkey, speak naturally, get polished text in any app.echogardenCross-platform speech toolset, used from the command-line or as a Node.js library. Includes a variety of engines for speech synthesis, speech recognition, forceFunCodecFunCodec is a research-oriented toolkit for audio quantization and downstream applications, such as text-to-speech synthesis, music generation et.al.speak-gptYour personal voice assistant based on OpenAI ChatGPT.speech-recognition-uk🇺🇦 Speech Recognition & Synthesis for UkrainianreverbOpen source inference code for Rev's modelklaamArabic speech recognition, classification and text-to-speech.awesome-ai-voiceList of open-source TTS, voice cloning, and music generation modelsbolnaEnd-to-end platform for building voice first multimodal agentsClickUiThe best way to use AI is on your own computer. Use local or paid API models, and ctrl+k to show/hide the chat UI. Experience the future of AI, and help build ireal-time-voice-translatorA desktop application that uses AI to translate voice between languages in real time, while preserving the speaker's tone and emotion.whisper-youtube🔉 Youtube Videos Transcription with OpenAI's WhisperVRCTVRCT(VRChat Chatbox Translator & Transcription)PreenCutAI-Powered Video Retrieval & Clipping ToolqvacOpen-source local AI SDK - run AI on-device with no cloud, no API keys. Supports GGUF, RAG, image, music, and video generation, speech-to-text, P2P inference, aawesome-audio-plazaDaily tracking of awesome audio papers, including music generation, zero-shot tts, asr, audio generationproject-ravenOpen-source AI meeting copilot - real-time transcription, echo cancellation, and AI assistance. Captures system audio + mic, cancels echo via WebRTC AEC3, transSynthalinguaSynthalingua - Real Time TranslationVoiceFlowLocal voice dictation and meeting recorder for Windows + Linux. Hold a hotkey to dictate, or record long-form meetings with system audio. Whisper transcription,awesome-russian-speechRussian speech technology linksai-avatar-system🎭 AI Avatar / digital human platform — upload a photo, clone a voice, talk to any face in real time with lip-sync video. Open-source, self-hosted. Claude · WhisMantellaMantella is a Skyrim and Fallout 4 mod which allows you to naturally speak to NPCs using a Speech-to-Text → LLMs → Text-to-Speech pipelineReazonSpeechMassive open Japanese speech corpusAudioToTextTranscribe and translate audio to text using Whisper and DeepL.xiaoniu小牛视频翻译 是一款支持本地视频翻译、字幕翻译和 YouTube 视频翻译下载的 AI 工具,集成自动语音识别与多语言翻译功能,助力创作者高效完成视频翻译,应用于视频本地化与视频出海场景。voicebook🗣️ A book and repo to get you started programming voice computing applications in Python (10 chapters and 200+ scripts).Stream-OmniStream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.parakeet-rsvery fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in RustVAF_2Aims to create a comprehensive voice toolkit for training, testing, and deploying speaker verification systems.wav2vec2-liveA live speech recognition using Facebooks wav2vec 2.0 model.Whisper-transcription_and_diarization-speaker-identification-How to use OpenAIs Whisper to transcribe and diarize audio filesbanini-tracker巴逆逆(8zz)反指標追蹤器 — Facebook 抓取 + 影片轉錄 + AI 分析 + 多平台推送(Telegram / Discord / LINE)tambourine-voiceYour personal voice interface for any app. Speak naturally and your words appear wherever your cursor is, with fully customizable AI voice dictation. Open soursubtitlerFree on-device web app for audio transcribing and rendering subtitlesizwiVoice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI compatible API.speechbrain.github.ioThe SpeechBrain project aims to build a novel speech toolkit fully based on PyTorch. With SpeechBrain users can easily create speech processing systems, rangingvoice_activity_detectionVoice Activity Detection based on Deep Learning & TensorFlowwhisper-obsidian-pluginSpeech-to-text in Obsidian using Whisperhms-ml-demoHMS ML Demo provides an example of integrating Huawei ML Kit service into applications. This example demonstrates how to integrate services provided by ML Kit, fcpx-auto-captions🎬 Auto Captions for Final Cut Pro Powered by OpenAI's Whisper ModelAwesome-Speaker-DiarizationSome comprehensive papers about speaker diarizationVectorDB-PluginProgram that lets you ask questions about your documents, audio, and video files.whisper-finetuneFine-tune and evaluate Whisper models for Automatic Speech Recognition (ASR) on custom datasets or datasets from huggingface.NovelDokushaAndroid web novel readernode-avFFmpeg bindings for Node.js. Features both low-level and high-level APIs, full hardware acceleration, TypeScript support, and modern async patternsLiveWhisperA nearly-live implementation of OpenAI's Whisper, using sounddevice. Requires existing Whisper install.Maix-SpeechMaix Speech AI lib, a fast and small speech lib running on embedded devices, including ASR, chat, TTS etc.onnx-asrA lightweight Python package for Automatic Speech Recognition using ONNX modelsyapFree, open source voice dictation for macOS. On-device transcription with Apple's Speech framework. No cloud/no API keys/no account.mlx-openai-serverA high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPI framework, it providesAIUIAIUI is a platform enabling seamless two-way verbal communication with AI.insanely-fast-whisper-apiAn API to transcribe audio with OpenAI's Whisper Large v3!WhisperHalluExperimental code: sound file preprocessing to optimize Whisper transcriptions without hallucinated textsLangHelperStriving to create a great Application with full functions of learning languages by ChatGPT, TTS, STT and other awesome AI models, supports talking, speaking asSpeech-to-Text-RussianПроект для распознавания речи на русском языке на основе pykaldi.speech_courseYSDA course in Speech Processing.whispercppPybind11 bindings for Whisper.cppWhisper-TikTokFrom AI tools to TikTok video creation using FFMPEG, Microsoft Edge read aloud and OpenAI Whisper modelqwen3-asrAll in one Qwen3-ASR Server, compatible with OpenAI APIwhisper-websiteSimple self-hosted web application, which can be used to convert audio to subtitles by OpenAI's Whisper modelhack-interviewAI-powered tool for real-time interview question transcription and response generation.openai-chat-api-workflow🎩 An Alfred 5 Workflow for using OpenAI Chat API to interact with GPT models 🤖💬 It also allows image generation/editing/understanding 🖼️, speech-to-text conversAliceAlice is a voice-first desktop AI assistant application built with Vue.js, Vite, and Electron. Advanced memory system, function calling, MCP support, optional fSayIt语音输入 + AI 润色,开源 Typeless 替代品。支持个人使用和团队/企业内部自部署。按下快捷键说话,文字自动输入到任何应用。UltraEval-AudioYour faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/TTS/Audio Codec/audio LvoiceaiSet of 📝 with 🔗 to help those building Voice AI agents 🎙️🤖whisper-nodeNode.js bindings for OpenAI's Whisper. (C++ CPU version by ggerganov)audio.cpp-webuiaudio.cpp with a full-task WebUI - pure C++ audio-model inference engine powered by ggml. TTS, ASR/STT, VAD, voice conversion, speaker diarization, music generaparakeet.cppUltra fast and portable Parakeet implementation for on-device inference in C++ using Axiom with MPS+Unified Memoryinput0Input0 — A macOS voice input tool: hold a hotkey to record, release to transcribe locally via STT, refine with LLM, and auto-paste into the active text field.ai-devicesAI Device Template Featuring Whisper, TTS, Groq, Llama3, OpenAI and moreyapsnapSnap any video URL or audio file into plaintext. No GPU. No cloud. One command.senkoVery fast, accurate speaker diarizationAriaARIA - AI Realtime Intelligent Audio | Universal real-time AI subtitles for WindowsSoulX-TranscriberAn end-to-end framework for multi-speaker transcription that jointly models who spoke, when, and what.watch-skillVideo understanding and self-verification for AI agents. Turn videos, streams, and agent screen recordings into searchable, timestamped evidence—then use THE LOasr-evaluationPython module for evaluating ASR hypotheses (e.g. word error rate, word recognition rate).sapphireShe's the AI agent you come home to.faster-whisper-webuia gradio webui for faster whispertemplate-tiktokGenerate TikTok-style captions with Whisper.cppT-oneT-one is a high-performance streaming ASR pipeline for Russian, specialized for the telephony domain.deepgram-js-sdkOfficial JavaScript SDK for Deepgram.praisesPraises is a text-to-speech tool that can help you read text easily.cobraOn-device voice activity detection (VAD) powered by deep learningtranscribeTranscribe is a real time transcription, conversation, Language learning platform. It provides live transcripts from microphone and speaker. It generates a suggspeechlibSpeechlib is a library that unifies speaker diarization, transcription and speaker recognition in a single pipeline to create transcripts for audio conversationDictateKeyboardA powerful Whisper AI keyboard for reliable speech transcriptionStage-WhisperThe main repo for Stage Whisper — a free, secure, and easy-to-use transcription app for journalists, powered by OpenAI's Whisper automatic speech recognition (Awhats-readerBrowse WhatsApp chat exports offline with AI-powered voice transcription. Privacy-first desktop/web app with bookmarks, search, and statistics. Built with SveltDeLiveSystem audio capture + multi-provider ASR + local-first AI review workspace. Floating live captions, 12 ASR backends, 60+ languages, AI summary/chat/mindmap, OpOn-Device-Speech-to-Speech-Conversational-AIThis is an on-CPU real-time conversational system for two-way speech communication with AI models, utilizing a continuous streaming architecture for fluid convegpt_servergpt_server是一个用于生产级部署LLMs、Embedding、Reranker、ASR、TTS、文生图、图片编辑和文生视频的开源框架。weeboA real-time speech-to-speech chatbot powered by Whisper Small, Llama 3.2, and Kokoro-82M.voice-assistant-whisper-chatgptThis repository will guide you to create your own Smart Virtual Assistant like Google Assistant using Open AI's ChatGPT, Whisper. The entire solution is createdLLMtunerFineTune LLMs in few lines of code (Text2Text, Text2Speech, Speech2Text)parakeetOpenAI Whisper-compatible ASR server using NVIDIA Parakeet TDT 0.6B (ONNX). Super fast CPU/GPU transcriptionsopenasrLocal-first speech-to-text: no cloud, no telemetry, fail-closed by design. One CLI, seven model families, signed model catalog, OpenAI-compatible local API.BreezeAppBreezeAPP 是一款為 Android 和 iOS 平台開發的純手機 AI 應用程式。從 App Store下載,即可在不連網的狀態下享受多項 AI 功能。源碼由聯發創新基地(MediaTek Research)提供。我們旨在推廣兩個概念: 人人都可以在自己的手機上自由選擇並運行不同的LLM - one is fAIThe definitive, open-source Swift framework for interfacing with generative AI.vid2cleantxtPython API & command-line tool to easily transcribe speech-based video files into clean textwhisper-auto-transcribeAuto transcribe tool based on whisperwtmBlazing fast whisper turbo for ASR (speech-to-text) tasksobs-cleanstreamCleanStream is an OBS plugin that uses AI to clean live audio streams from unwanted words and utterancesShushShush is an app that deploys a WhisperV3 model with Flash Attention v2 on Modal and makes requests to it via a NextJS appobsidian-transcriptionObsidian plugin to create high-quality transcriptions from markdown linked audio filesinsights-lm-local-packageOpen-source, fully private and local alternative to NotebookLM. Chat with your documents, generate audio summaries, and ground AI in your own sources—built withtggeneratorGenerador de logotipos de eSports por IA (con fines académicos durante el evento Tenerife GG)milaMila — native macOS local transcription app (whisper.cpp) with optional speaker diarization. Apache-2.0.vox-boxA text-to-speech and speech-to-text server compatible with the OpenAI API, supporting Whisper, FunASR, Bark, and CosyVoice backends.nodejs-whisperNodeJS Bindings for Whisper - the CPU version of OpenAI's Whisper, as initially crafted in C++ by ggerganov.Local-Multimodal-AI-ChatSelf-hostable multimodal chat with local LLMs (Ollama/OpenAI): PDF RAG, image chat, and Whisper voice, Streamlit + Docker.wyoming_openaiOpenAI-Compatible Proxy Middleware for the Wyoming Protocolthonburian-whisperThonburian Whisper: Open models for fine-tuned Whisper in Thai. Try our demo on Huggingface space:whisper-live-transcriptionLive-Transcription (STT) with Whisper PoChumlaOpen-source AI meeting notes for Mac. Records mic + system audio with no bot, transcribes on-device or via OpenAI / Deepgram / Groq, identifies speakers offlineautoshortsAutomatically generate viral-ready vertical short clips from long-form gameplay footage using AI-powered scene analysis, GPU-accelerated rendering, and optionalwhisper_real_time_translationThe subtitles and translations are generated in real-time and displayed as pop-ups.typewhisper-winTypeWhisper for Windows - Local speech-to-text with translationspeedofsoundVoice typing for the Linux desktop.evaA New End-to-end Framework for Evaluating Voice AgentsrustfstRust re-implementation of OpenFST - library for constructing, combining, optimizing, and searching weighted finite-state transducers (FSTs). A Python binding isFS-EENDThe official Pytorch implementation of "Frame-wise streaming end-to-end speaker diarization with non-autoregressive self-attention-based attractors". [ICASSP 20earshotRidiculously fast & accurate streaming voice activity detectionmatrix-live-diarizerLocal-first meeting transcription — audio & transcripts never leave your machine. Live captions + upload diarization + voice matching.whisperX-FastAPIFastAPI service on top of WhisperXAidgetAi edge toolbox,专门面向边端设备尤其是嵌入式RTOS平台,AI模型部署工具链,包括模型推理引擎和模型压缩工具BlahSTInput text from speech in any Linux window, the lean, fast and accurate way, using whisper.cpp OFFLINE. Speak with local LLMs via llama.cpp.meetevalMeetEval - A meeting transcription evaluation toolkitcv-datasetMetadata and versioning details for the Common Voice datasetviet-asrVietASR - Vietnamese Automatic Speech Recognitionsova-asrSOVA ASR (Automatic Speech Recognition)WhisperTimeSyncSynchronize Whisper's timestamps over an existing accurate transcriptionspinoramaA library to display and compare speakers and headphones measurements.MNNServerA third-party MNN server supporting external calls, embedding model, TTS model and ASR model features.一个支持外部调用、向量模型、文字转语音模型和语音识别模型特性的第三方MNN服务器speechTAn opensource speech-to-text software written in tensorflowQSmartAssistant一个模块化,全过程可离线,低占用率的对话机器人/智能音箱livecaptionReal-time on-device speech transcription + translation for macOS (Apple Silicon). Streaming ASR, speaker diarization, and live English→Chinese translation on Apkroko-onnxKroko ASR - Speech-to-textsimple_diarizerSimplified diarization pipeline using some pretrained models - audio file to diarized segments in a few lines of codeLiteASR[EMNLP Main '25] LiteASR: Efficient Automatic Speech Recognition with Low-Rank ApproximationFunSpeech开箱即用的本地私有化部署语音服务,快速搭建Qwen3ASR/FunASR与Qwen3TTS/CosyVoice后端rVADfastThis is the Python library for an unsupervised, fast method for robust voice activity detection (rVAD), as in the paper rVAD: An Unsupervised Segment-Based Robueuanka本地完整部署ASR(K2)-NLP(Rasa,Spacy)-LLM(Chatglm2)-TTS(Vits)ASR-Wav2vec-Finetune:zap: Finetune Wa2vec 2.0 For Speech Recognitionwhisply💬 Fast, cross-platform CLI and GUI for batch transcription, translation, speaker annotation and subtitle generation using OpenAI’s Whisper on CPU, Nvidia GPU anwhisperVideoFind out who said what in the video.SqueezeformerPyTorch implementation of "Squeezeformer: An Efficient Transformer for Automatic Speech Recognition" (NeurIPS 2022)tafrighتفريغ النصوص وإنشاء ملفات SRT و VTT باستخدام نماذج Whisper وتقنية wit.ai.PunctuationModel中文标点符号模型,可以给文本添加标点符号。hypercheap-voiceAIThe most cost-effective, highest performance AI voice agent possible todayGPVRepository for our Interspeech2020 general-purpose voice activity detection (GPVAD) paperspeech-androidOn-device speech SDK for Android — ASR, TTS, VAD, and noise cancellation powered by ONNX Runtime with Qualcomm NNAPI accelerationrVADMatlab and Python libraries for an unsupervised method for robust voice activity detection (rVAD), as in the paper rVAD: An Unsupervised Segment-Based Robust Vovosk-asteriskSpeech Recognition in Asterisk with Vosk ServeritspIntroduction to Speech Processingmeeting-transcriberOn-device meeting transcriber for macOS — auto-records Teams/Zoom/Webex, transcribes & separates speakers locally. No cloud. Open-source alternative to Otter/Grsaa-sdkAddressee detection for voice agents: device-directed speech detection that runs before STT, so background speech, side conversations, and the agent's own TTS ediarizeSpeaker diarization for Python — "who spoke when?" CPU-only, no API keys, Apache 2.0. ~10.8% DER on VoxConverse, 8x faster than real-time.zanshinA novel media player that allows you to navigate by speakerspeakrsSpeaker diarization in Rust. 312–912x realtime on Apple Silicon, 50–121x on CUDA. Matches pyannote accuracy.IS2023-powerset-diarizationOfficial repository for the "Powerset multi-class cross entropy loss for neural speaker diarization" paper published in Interspeech 2023.TargetDiarizationMulti-speaker separation, identification, diarization ALL-IN-ONE. It can isolate the target speaker from a conversation audio and do ASR.whisper_rosSpeech-to-Text based on SileroVAD + whisper.cpp (GGML Whisper) for ROS 2dub-studioFree offline AI video dubbing studio for Windows — voice cloning, translation, subtitles & on-screen-text localization. 100% local, one native .exe, zero PythonDatadriven-GPVADThe codebase for Data-driven general-purpose voice activity detection.docker-whisperDocker image for a self-hosted Whisper speech-to-text server with speaker diarization and OpenAI-compatible transcription and translation APIs. Powered by fastevoice_gender_detection♂️♀️ Detect a person's gender from a voice file (90.7% +/- 1.3% accuracy).D-TDNNPyTorch implementation of Densely Connected Time Delay Neural NetworkvoxsegA python library for voice activity detection (VAD) for speech/non-speech segmentation.OpenTranscribeSelf-hosted AI-powered transcription platform with speaker diarization, search, and collaboration features. Built with Svelte, FastAPI, and Docker for easy deplspeech-condenserA tool for summarizing dialogues from videos or audioCallyticsCallytics is an advanced call analytics solution that leverages speech recognition and large language models (LLMs) technologies to analyze phone conversations falconOn-device speaker diarization powered by deep learningawesome-vadA curated list of awesome voice activity detectiongryannoteProvide Gradio custom components to make the diarization-based audio labeling process easier and faster.wyoming-voice-matchA Wyoming protocol ASR proxy that verifies speaker identity and isolates voice commands from background noise before forwarding audio to a downstream speech-to-speech-coreOn-device VAD / streaming STT / TTS / diarization in C++17 (ONNX + LiteRT) with a voice-agent pipeline. Linux, Windows, Android.gemini-transcribeTranscribe audio and video files with speaker diarization and logically grouped timestamps using Gemini FlashSimpleDERA lightweight library to compute Diarization Error Rate (DER).ai-clips-makerAI-powered tool to turn long videos into short, viral-ready clips. Combines transcription, speaker diarization, scene detection & 9:16 resizing — perfect for crSimpleDiarizationSimple diarization modelvoicetagSpeaker identification powered by pyannote and resemblyzerMLC-SLM-BaselineThe project is associated with the recently-launched INTERSPEECH 2025 Workshop on Multilingual Conversational Speech Language Model (MLC-SLM) to provide particRE-VERBspeaker diarization system using an LSTMrttm-viewerApplication for viewing Rich Transcription Time Marked (RTTM) files in an interactive waysepia-web-audioCreate modular, cross-browser, web audio pipelines to record and process audio in background threads. Comes with modules for VAD, ASR, resampling and much more.mamba-diarizationOfficial repository for Mamba-based Segmentation Model for Speaker DiarizationNAMO-Turn-Detector-v1High-performance, semantic turn detection for conversational AIvoice-agents-from-scratchFrom-scratch voice agents in Python: end-to-end speech pipelines, runnable chapters, and a small shared library. Local models, explicit streaming behavior.WhisperSegCode for ICASSP 2024 paper WhisperSeg: Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity DetectionspectraSpectra extraction tutorials based on torch and torchaudio.ai-speech-engineer-roadmap🐌 A curated roadmap based on my 6 years of experience form zero to become a skilled AI Speech Engineer. This roadmap covers everything from fundamentals to cuttvad[Tiny VAD] SG-VAD: Stochastic Gates Based Speech Activity Detection