🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
- Updated
Jul 26, 2026 - Python
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
SRS is a simple, high-efficiency, real-time media server supporting RTMP, WebRTC, HLS, HTTP-FLV, HTTP-TS, SRT, MPEG-DASH, and GB28181, with codec support for H.264, H.265, AV1, VP9, AAC, Opus, and G.711.
GUI for a Vocal Remover that uses Deep Neural Networks.
Javascript audio library for the modern web.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
BlackHole is a modern macOS audio loopback driver that allows applications to pass audio to other applications with zero additional latency.
Background Music, a macOS audio utility: automatically pause your music, set individual apps' volumes and record system audio.
The free and privacy-friendly screen recorder with no limits 🎥
💿 Free software that works great, and also happens to be open-source Python.
Music streaming solution that works.
Pure Go implementation of the WebRTC API
A hands-on introduction to video technology: image, video, codec (av1, vp9, h265) and more (ffmpeg encoding). Translations: 🇺🇸 🇨🇳 🇯🇵 🇮🇹 🇰🇷 🇷🇺 🇧🇷 🇪🇸
Simple and Fast Multimedia Library
Code. Music. Live.
A PyTorch-based Speech Toolkit