VoxCPM
VoxCPM is a tokenizer-free Text-to-Speech system that directly generates continuous speech representations via an end-to-end diffusion autoregressive architecture, bypassing discrete tokenization to achieve highly natural and expressive synthesis. VoxCPM2 is the latest major release, a 2B parameter model trained on over 2 million hours of multilingual speech data, now supporting 30 languages, Voice Design, Controllable Voice Cloning, and 48kHz studio-quality audio output. Built on a MiniCPM-4 backbone. 30-Language Multilingual, Input text in any of the 30 supported languages and synthesize directly, no language tag needed Voice Design, Create a brand-new voice from a natural-language description alone (gender, age, tone, emotion, pace …), no reference audio required Controllable Cloning, Clone any voice from a short reference clip, with optional style guidance to steer emotion, pace, and expression while preserving the original timbre Ultimate Cloning, Reproduce every vocal nuance: provide both reference audio and its transcript, and the model continues seamlessly from the reference, faithfully preserving every vocal detail, timbre, rhythm, emotion, and style (same as VoxCPM1.5) 48kHz High-Quality Audio, Accepts 16kHz reference audio and directly outputs 48kHz studio-quality audio via AudioVAE V2's asymmetric encode/decode design, with built-in super-resolution, no external upsampler needed Context-Aware Synthesis, Automatically infers appropriate prosody and expressiveness from text content Real-Time Streaming, RTF as low as ~0.3 on NVIDIA RTX 4090, and ~0.13 accelerated by Nano-vLLM or vLLM-Omni, official vLLM omni-modal serving for VoxCPM2 with PagedAttention and an OpenAI-compatible API Fully Open-Source & Commercial-Ready, Weights and code released under the Apache-2.0 license, free for commercial use Chinese Dialect: 四川话, 粤语, 吴语, 东北话, 河南话, 陕西话, 山东话, 天津话, 闽南话 [2026.04] We release VoxCPM2, 2B, 30 languages, Voice Design & Controllable Voice Cloning, 48kHz audio output! Weights Docs Playground Technical Report [2025.12] Open-source VoxCPM1.5 weights with SFT & LoRA fine-tuning. ( #1 GitHub Trending) [2025.09] Release VoxCPM Technical Report. [2025.09] Open-source VoxCPM-0.5B weights ( #1 HuggingFace Trending)
View VoxCPM on GitHub
VoxCPM is a tokenizer-free Text-to-Speech system that directly generates continuous speech representations via an end-to-end diffusion autoregressive architecture, bypassing discrete tokenization to achieve highly natural and expressive synthesis.
VoxCPM2 is the latest major release, a 2B parameter model trained on over 2 million hours of multilingual speech data, now supporting 30 languages, Voice Design, Controllable Voice Cloning, and 48kHz studio-quality audio output. Built on a MiniCPM-4 backbone. 30-Language Multilingual, Input text in any of the 30 supported languages and synthesize directly, no language tag needed Voice Design, Create a brand-new voice from a natural-language description alone (gender, age, tone, emotion, pace …), no reference audio required Controllable Cloning, Clone any voice from a short reference clip, with optional style guidance to steer emotion, pace, and expression while preserving the original timbre Ultimate Cloning, Reproduce every vocal nuance: provide both reference audio and its transcript, and the model continues seamlessly from the reference, faithfully preserving every vocal detail, timbre, rhythm, emotion, and style (same as VoxCPM1.5) 48kHz High-Quality Audio, Accepts 16kHz reference audio and directly outputs 48kHz studio-quality audio via AudioVAE V2's asymmetric encode/decode design, with built-in super-resolution, no external upsampler needed Context-Aware Synthesis, Automatically infers appropriate prosody and expressiveness from text content Real-Time Streaming, RTF as low as ~0.3 on NVIDIA RTX 4090, and ~0.13 accelerated by Nano-vLLM or vLLM-Omni, official vLLM omni-modal serving for VoxCPM2 with PagedAttention and an OpenAI-compatible API Fully Open-Source & Commercial-Ready, Weights and code released under the Apache-2.0 license, free for commercial use
Chinese Dialect: 四川话, 粤语, 吴语, 东北话, 河南话, 陕西话, 山东话, 天津话, 闽南话 [2026.04] We release VoxCPM2, 2B, 30 languages, Voice Design & Controllable Voice Cloning, 48kHz audio output! Weights Docs Playground Technical Report [2025.12] Open-source VoxCPM1.5 weights with SFT & LoRA fine-tuning. ( #1 GitHub Trending) [2025.09] Release VoxCPM Technical Report. [2025.09] Open-source VoxCPM-0.5B weights ( #1 HuggingFace Trending)
VoxCPM at a glance
| Stars | 38k |
|---|---|
| Forks | 4.3k |
| Language | Python |
| License | Apache-2.0 |
| Last update | 2026-09-02 |
| Contributors | 32 |
How to install VoxCPM
bash pip install voxcpm
Where VoxCPM is listed
Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.
GitHub repos like this
More repo topics
More free tools
Related MCP servers & CLIs
The most actionable AI newsletter for founders
Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.
No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.
P.S. Sign up now to get free access to my ultimate AI tools guide for creators.






































