[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
Categories & Topics
Repository Stats
Related Tools
audio.cpp
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
FluidAudio
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
Fun-ASR
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
KittenTTS
KittenTTS enables users to generate lifelike speech from text, making it easier to create engaging audio content. With a focus on accessibility, it allows anyone to transform written information into spoken words effortlessly.
LuxTTS
LuxTTS is a voice cloning model that allows for high-quality text-to-speech conversion, achieving speeds up to 150 times faster than real-time. It offers an efficient solution for creating realistic voiceovers quickly, making it valuable for various applications in media and communication.
MOSS-TTS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.