Navid MoazzezNavid Moazzez

VoxCPM

VoxCPM is a tokenizer-free Text-to-Speech system that directly generates continuous speech representations via an end-to-end diffusion autoregressive architecture, bypassing discrete tokenization to achieve highly natural and expressive synthesis. VoxCPM2 is the latest major release, a 2B parameter model trained on over 2 million hours of multilingual speech data, now supporting 30 languages, Voice Design, Controllable Voice Cloning, and 48kHz studio-quality audio output. Built on a MiniCPM-4 backbone. 30-Language Multilingual, Input text in any of the 30 supported languages and synthesize directly, no language tag needed Voice Design, Create a brand-new voice from a natural-language description alone (gender, age, tone, emotion, pace …), no reference audio required Controllable Cloning, Clone any voice from a short reference clip, with optional style guidance to steer emotion, pace, and expression while preserving the original timbre Ultimate Cloning, Reproduce every vocal nuance: provide both reference audio and its transcript, and the model continues seamlessly from the reference, faithfully preserving every vocal detail, timbre, rhythm, emotion, and style (same as VoxCPM1.5) 48kHz High-Quality Audio, Accepts 16kHz reference audio and directly outputs 48kHz studio-quality audio via AudioVAE V2's asymmetric encode/decode design, with built-in super-resolution, no external upsampler needed Context-Aware Synthesis, Automatically infers appropriate prosody and expressiveness from text content Real-Time Streaming, RTF as low as ~0.3 on NVIDIA RTX 4090, and ~0.13 accelerated by Nano-vLLM or vLLM-Omni, official vLLM omni-modal serving for VoxCPM2 with PagedAttention and an OpenAI-compatible API Fully Open-Source & Commercial-Ready, Weights and code released under the Apache-2.0 license, free for commercial use Chinese Dialect: 四川话, 粤语, 吴语, 东北话, 河南话, 陕西话, 山东话, 天津话, 闽南话 [2026.04] We release VoxCPM2, 2B, 30 languages, Voice Design & Controllable Voice Cloning, 48kHz audio output! Weights Docs Playground Technical Report [2025.12] Open-source VoxCPM1.5 weights with SFT & LoRA fine-tuning. ( #1 GitHub Trending) [2025.09] Release VoxCPM Technical Report. [2025.09] Open-source VoxCPM-0.5B weights ( #1 HuggingFace Trending)

View VoxCPM on GitHub
Navid Moazzezby Navid Moazzez·Updated Sept 30, 2026·2 min read
VoxCPM

VoxCPM is a tokenizer-free Text-to-Speech system that directly generates continuous speech representations via an end-to-end diffusion autoregressive architecture, bypassing discrete tokenization to achieve highly natural and expressive synthesis.

VoxCPM2 is the latest major release, a 2B parameter model trained on over 2 million hours of multilingual speech data, now supporting 30 languages, Voice Design, Controllable Voice Cloning, and 48kHz studio-quality audio output. Built on a MiniCPM-4 backbone. 30-Language Multilingual, Input text in any of the 30 supported languages and synthesize directly, no language tag needed Voice Design, Create a brand-new voice from a natural-language description alone (gender, age, tone, emotion, pace …), no reference audio required Controllable Cloning, Clone any voice from a short reference clip, with optional style guidance to steer emotion, pace, and expression while preserving the original timbre Ultimate Cloning, Reproduce every vocal nuance: provide both reference audio and its transcript, and the model continues seamlessly from the reference, faithfully preserving every vocal detail, timbre, rhythm, emotion, and style (same as VoxCPM1.5) 48kHz High-Quality Audio, Accepts 16kHz reference audio and directly outputs 48kHz studio-quality audio via AudioVAE V2's asymmetric encode/decode design, with built-in super-resolution, no external upsampler needed Context-Aware Synthesis, Automatically infers appropriate prosody and expressiveness from text content Real-Time Streaming, RTF as low as ~0.3 on NVIDIA RTX 4090, and ~0.13 accelerated by Nano-vLLM or vLLM-Omni, official vLLM omni-modal serving for VoxCPM2 with PagedAttention and an OpenAI-compatible API Fully Open-Source & Commercial-Ready, Weights and code released under the Apache-2.0 license, free for commercial use

Chinese Dialect: 四川话, 粤语, 吴语, 东北话, 河南话, 陕西话, 山东话, 天津话, 闽南话 [2026.04] We release VoxCPM2, 2B, 30 languages, Voice Design & Controllable Voice Cloning, 48kHz audio output! Weights Docs Playground Technical Report [2025.12] Open-source VoxCPM1.5 weights with SFT & LoRA fine-tuning. ( #1 GitHub Trending) [2025.09] Release VoxCPM Technical Report. [2025.09] Open-source VoxCPM-0.5B weights ( #1 HuggingFace Trending)

VoxCPM at a glance

Stars38k
Forks4.3k
LanguagePython
LicenseApache-2.0
Last update2026-09-02
Contributors32

How to install VoxCPM

bash pip install voxcpm 

Where VoxCPM is listed

Navid Moazzez

AI business strategist & AI OS builder

Navid Moazzez helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life.

Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.

GitHub repos like this

More free tools

Related MCP servers & CLIs

Free AI newsletter

The most actionable AI newsletter for founders

Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.

No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.

P.S. Sign up now to get free access to my ultimate AI tools guide for creators.

Loved by 10,000+ readers