Audio HubAudio AI

Audio AI

134 tools and products, each with a live profile and alternatives list.

essentiaC++ library for audio and music analysis, description and synthesis, including Python bindingsSpleeterGuiWindows desktop front end for Spleeter - AI source separationfmaFMA: A Dataset For Music AnalysisasteroidThe PyTorch-based audio source separation toolkit for researchersomnizartOmniscient Mozart, being able to transcribe everything in the music, including vocal, drum, chord, beat, instruments, and more.madmomPython audio and music signal processing librarymeydaAudio feature extraction for JavaScript.Semi-supervised-learningA Unified Semi-Supervised Learning Codebase (NeurIPS'22)astCode for the Interspeech 2021 paper "AST: Audio Spectrogram Transformer".crepeCREPE: A Convolutional REpresentation for Pitch Estimation -- pre-trained model (ICASSP 2018)musicinformationretrieval.comInstructional notebooks on music information retrieval.voicefilterUnofficial PyTorch implementation of Google AI's VoiceFilter systemessentia.jsJavaScript library for music/audio analysis and processing powered by Essentia WebAssemblyawesome-audio-visualA curated list of different papers and datasets in various areas of audio-visual processingConv-TasNetA PyTorch implementation of Conv-TasNet described in "TasNet: Surpassing Ideal Time-Frequency Masking for Speech Separation" with Permutation Invariant TrainingSOMESOME: Singing-Oriented MIDI Extractor.pitch-detectionautocorrelation-based O(NlogN) pitch detectionAudio-ClassificationCode for YouTube series: Deep Learning for Audio ClassificationmsafMusic Structure Analysis Frameworkspleeter-webSelf-hostable web app for isolating the vocal, accompaniment, bass, and drums of any song. Supports Spleeter, Demucs, BS-RoFormer. Started in 2019.Urban-Sound-ClassificationUrban sound classification using Deep LearningexamplesAnalyze the unstructured data with Towhee, such as reverse image search, reverse video search, audio classification, question and answer systems, molecular searHTS-Audio-TransformerThe official code repo of "HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection"spafe:sound: spafe: Simplified Python Audio Features ExtractionOlafOlaf: Overly Lightweight Acoustic Fingerprinting is a portable acoustic fingerprinting system.GistA C++ Library for Audio Analysismtg-jamendo-datasetMetadata, scripts and baselines for the MTG-Jamendo datasetOMR-DatasetsCollection of datasets used for Optical Music RecognitionPaSSTEfficient Training of Audio Transformers with PatchouttutorialTutorial covering Open Source tools for Source Separation.llarkCode for the paper "LLark: A Multimodal Instruction-Following Language Model for Music" by Josh Gardner, Simon Durand, Daniel Stoller, and Rachel Bittner.spotifyrR wrapper for Spotify's Web APIbeat_thisAccurate and general beat trackerCLMROfficial PyTorch implementation of Contrastive Learning of Musical Representationshover:speedboat: Label data at scale. Fun and precision included.pb_bssCollection of EM algorithms for blind source separation of audio signalsmultimodal-ml-musicList of academic resources on Multimodal ML for Musicpitch-detectionA collection of algorithms to determine the pitch of a sound sample.DatasetsDatasets of MusicBrainz, Tidal, Spotify, DeezerMusic-Emotion-RecognitionA Machine Learning Approach of Emotional Modelmdx-netKUIELAB-MDX-Net got the 2nd place on the Leaderboard A and the 3rd place on the Leaderboard B in the MDX-Challenge ISMIR 2021EAT[IJCAI 2024] EAT: Self-Supervised Pre-Training with Efficient Audio Transformerneural-audio-fpOfficial implementation of Neural Audio Fingerprint (ICASSP 2021)BandSplitRNN-PyTorchUnofficial PyTorch implementation of Music Source Separation with Band-split RNNMARBLEState-of-the-art pretrained music models for training, evaluation, inferenceswift-f0Fast and accurate fundamental frequency (F0) detector using convolutional neural networksAudio-Mamba-AuMOfficial Implementation of the work "Audio Mamba: Bidirectional State Space Model for Audio Representation Learning"TuneNNA transformer-based network model for pitch detectionvocalsoundDataset and baseline code for the VocalSound dataset (ICASSP2022).bliss-rsA song analysis library for making playlistsautochordAutomatic Chord Recognition tools - ISMIR2021 Late-Breaking Demo presentationMAX-Audio-ClassifierIdentify sounds in short audio clipspslaCode for the TASLP paper "PSLA: Improving Audio Tagging With Pretraining, Sampling, Labeling, and Aggregation".audiosslA library built for easier audio self-supervised training, downstream tasks evaluationssamba[SLT'24] The official implementation of SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Modelchord-detectionApp for Chord Sequence Detectionuvr-mdx-inferUltimate Vocal Remover Inference CLImuscallOfficial implementation of "Contrastive Audio-Language Learning for Music" (ISMIR 2022)dcase2020_task2_baselineDCASE2020 Challenge Task 2 baseline systemACA-SlidesSlides and Code for "An Introduction to Audio Content Analysis," also taught at Georgia Tech as MUSI-6201. This introductory course on Music Information RetrievbeetcampBandcamp autotagger source for beets (https://beets.io)graphmuseA Graph Deep Learning Library for Music.mad-twinnetThe code for the MaD TwinNet. Demo page:DTTNet-PytorchAn official implementation of the ICASSP 2024 paper: Dual-Path TFC-TDF UNet for Music Source SeparationchocoChoCo: the Chord CorpusESC-CNN-microcontrollerEnvironmental Sound Classification on Microcontrollers using Convolutional Neural Networksquery-banditBanquet: A Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four StemsfastF0NlsC++ and MATLAB code for fast and accurate fundamental frequency estimationAudioClassification-PaddlePaddle基于PaddlePaddle实现的音频分类,支持EcapaTdnn、PANNS、TDNN、Res2Net、ResNetSE等各种模型,还有多种预处理方法VirtualConductor🎶 Music-Driven Conducting Motion Generation (IEEE ICME'21 Best Demo)whatbpm💓 Today's Trending Values for EDM Productionsdx23Sound Demixing Challenge 2023SymbTrTurkish Makam Music Symbolic Data Collectionjazznetjazznet dataset of piano patterns for music audio machine learning researchAwesome-Music-Recommendation-DatasetsAwesome Datasets for Music RecommendationConditioned-Source-Separation-LaSAFTA PyTorch implementation of the paper: "LaSAFT: Latent Source Attentive Frequency Transformation for Conditioned Source Separation" (ICASSP 2021)ACA-CodeMatlab scripts accompanying the book "An Introduction to Audio Content Analysis" (www.AudioContentAnalysis.org)AI-Generated-Content-DetectionMulti-modal AI-generated content detection: image, video, and audio. Benchmarks, training code (DINOv2, DINOv3, ReStraV, BreathNet), and evaluation pipeline formuscapsSource code for "MusCaps: Generating Captions for Music Audio" (IJCNN 2021)demucs_lightningDemucs Lightning: A PyTorch lightning version of Demucs with Hydra and Tensorboard featuresaudio_source_separationAn implementation of audio source separation tools.libACAC++ code accompanying the book "An Introduction to Audio Content Analysis" (www.AudioContentAnalysis.org)Audio_Classification_using_LSTMClassification of Urban Sound Audio Dataset using LSTM-based model.SemiReward[ICLR 2024] SemiReward: A General Reward Model for Semi-supervised LearningmeicoA converter framework with support for MEI, MSM, MPM, MIDI, WAV, MP3, chroma, and XSLTflutter_tflite_audioAudio classification Tflite package for flutter (iOS & Android). Can support Google Teachable Machine modelsflutter-fftFlutter pitch detection/audio processing plugin, personalized for my guitar tuner application.ml-audio-classifier-example-for-picoML Audio Classifier Example for Pico 🔊🔥🔔ismir2018-revisiting-svdRevisiting Singing Voice Detection : a Quantitative Review and the Future OutlookdechorderAutomatic chord recognition application powered by machine learningLos-Angeles-MIDI-DatasetSOTA kilo-scale MIDI dataset for MIR and Music AI purposesda-tacosA Dataset for Cover Song Identification and UnderstandinghfapigoUnofficial (Golang) Go bindings for the Hugging Face Inference APISCNet-PyTorchUnofficial PyTorch implementation of "SCNet: Sparse Compression Network for Music Source Separation"musicntwrkNetwork Analysis of Generalized Musical SpacesbabycatAn audio manipulation library for Rust, Python, WebAssembly, and C.speakerboxSpeakerbox: Fine-tune Audio Transformers for speaker identification.music-to-midi音乐转MIDI转换器 - Convert audio to multi-track MIDI with lyrics embeddingbeets-beatport4Beatport API v4 compatible beets pluginarrangerOfficial Implementation of "Towards Automatic Instrumentation by Learning to Separate Parts in Symbolic Multitrack Music" (ISMIR 2021)voc2vecThis repository contains the code for the paper "voc2vec: A Foundation Model for Non-Verbal Vocalization", accepted at ICASSP 2025.sampleCNN-pytorchPytorch implementation of "Sample-level Deep Convolutional Neural Networks for Music Auto-tagging Using Raw Waveforms"scarlethyperspectral galaxy modeling and deblendinguavmCode for the IEEE Signal Processing Letters 2022 paper "UAVM: Towards Unifying Audio and Visual Models".o-m_beatmap_trainerTraining pipeline for an osu!mania 7k next-event model, designed as the upstream predictor for audio-driven 7k map generation systems.mxnet-audioImplementation of music genre classification, audio-to-vec, song recommender, and music search in mxnetreact-native-pitchyA real-time pitch detection library for React Native.sonics[ICLR 2025] SONICS: Synthetic Or Not - Identifying Counterfeit Songsjdcnet-pytorchpytorch implementation of JDCNet, singing voice detection and classification networkTunaPitch detection & utils.FastTuneA tuner for guitar, ukulele, bass, banjo, mandolin, violin and etc.Vision_Audio_and_Multimodal_ProjectsThis repository includes all computer vision, audio, document AI, and multimodal projects.librosa.cppC++17 port of librosa with wasm and SPM package. Done with agents. YMMV.AudioPitchEstimatorForUnityA simple real-time pitch estimator for UnityVAE-BSSUnsupervised blind source separation of mixed images and sounds with variational auto-encoders.aiSFXRepresentation Learning for the Automatic Indexing of Sound Effects Libraries (ISMIR 2022): Deep audio embeddings pre-trained on UCS & Non-UCS-compliant datasetHPSSHarmonic/Percussive Sound SeparationRizzRiffMixed Reality Guitar App on the Meta Quest platformdeepperformerDeep Performer: Score-to-audio music performance synthesisArduino-FrequencyDetectorFast audio frequency detector without fft for plain Arduino and Attiny85. Whistle switch example included.MMSP2021-Audio2ScoreAlignmentAudio-to-Score Alignment Using Deep Automatic Music Transcriptionsdx23-aimlessSource Separation training codebase for the Sound Demixing Challenge 2023.musicbucket-bot-oldA Telegram bot that helps chat users sharing and keeping track musicaudio-classification-pytorchIn this project, several approaches for training/finetuning an audio gender recognition is provided. The code can simply be used for any other audio classificatllm-tseTyping to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction (LLM-TSE)vocalA vocal source separationMIDImikePitch Detection on Arduino using Autocorrelationireal-readerA Node JS module to read music files from iRealPro.SiTraNoA MATLAB app for sines-transient-noise decomposition of audio signals.ScorePerformerScorePerformer: Expressive Piano Performance Rendering with Fine-Grained Control (ISMIR 2023)sourceA Freesound Community SamplerTONetThe official implementation of "TONet: Tone-Octave Network for Singing Melody Extraction from Polyphonic Music"bpm-detectorA Python tool for automatic detection of BPM (tempo) and musical key from audio files.WebSpeechAnalyzerJS speech analyzer for fast speech analysis and labeling