VoxCPM enables the generation of natural-sounding speech and lifelike voice cloning without the need for traditional tokenizers. It leverages advanced deep learning techniques to create context-aware audio, making it ideal for applications in text-to-speech technology.
Categories & Topics
Repository Stats
Related Tools
GPA
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
KittenTTS
KittenTTS enables users to generate lifelike speech from text, making it easier to create engaging audio content. With a focus on accessibility, it allows anyone to transform written information into spoken words effortlessly.
LuxTTS
LuxTTS is a voice cloning model that allows for high-quality text-to-speech conversion, achieving speeds up to 150 times faster than real-time. It offers an efficient solution for creating realistic voiceovers quickly, making it valuable for various applications in media and communication.
MOSS-TTS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
neutts
On-device TTS model by Neuphonic
OmniVoice
High-Quality Voice Cloning TTS for 600+ Languages