Navid MoazzezNavid Moazzez

AirLLM

AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card, without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB, DeepSeek-V3 (671B) on ~12GB, and Kimi K3 (2.8T), the largest open-source model released to date, on under 4GB, because sparse MoE models stream one expert at a time rather than a whole layer. [2026/07] Kimi K3 (2.8T) support: the largest open-source model runs on a single card in 3.72GB of VRAM, measured end to end on one RTX 6000 Ada. Per-expert streaming loads only the experts a token actually routes to. K3 brings three requirements of its own: pip install compressed-tensors flash-attn (its model code mandates flash attention regardless of what you request), a CUDA 12 build of torch, since no prebuilt flash-attn wheel exists for CUDA 13 yet, and transformers 4.56.x, as its remote code does not load on 5.x. [2026/06] v3.0: FP8 model support + the latest models. Run DeepSeek-V3 (671B) on ~12GB and Qwen3-235B on ~3GB, plus Qwen3, Llama 3.x/4, DeepSeek V2/V3, Phi-4, Gemma and more, all through a single AutoModel.

View AirLLM on GitHub
Navid Moazzezby Navid Moazzez·Updated Sept 30, 2026·1 min read
AirLLM

AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card, without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB, DeepSeek-V3 (671B) on ~12GB, and Kimi K3 (2.8T), the largest open-source model released to date, on under 4GB, because sparse MoE models stream one expert at a time rather than a whole layer.

[2026/07] Kimi K3 (2.8T) support: the largest open-source model runs on a single card in 3.72GB of VRAM, measured end to end on one RTX 6000 Ada. Per-expert streaming loads only the experts a token actually routes to. K3 brings three requirements of its own: pip install compressed-tensors flash-attn (its model code mandates flash attention regardless of what you request), a CUDA 12 build of torch, since no prebuilt flash-attn wheel exists for CUDA 13 yet, and transformers 4.56.x, as its remote code does not load on 5.x.

[2026/06] v3.0: FP8 model support + the latest models. Run DeepSeek-V3 (671B) on ~12GB and Qwen3-235B on ~3GB, plus Qwen3, Llama 3.x/4, DeepSeek V2/V3, Phi-4, Gemma and more, all through a single AutoModel.

AirLLM at a glance

Stars35k
Forks3.7k
LanguageJupyter Notebook
LicenseApache-2.0
Last update2026-09-28
Contributors10

How to install AirLLM

bash pip install airllm 

Where AirLLM is listed

Navid Moazzez

AI business strategist & AI OS builder

Navid Moazzez helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life.

Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.

GitHub repos like this

More free tools

Related MCP servers & CLIs

Free AI newsletter

The most actionable AI newsletter for founders

Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.

No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.

P.S. Sign up now to get free access to my ultimate AI tools guide for creators.

Loved by 10,000+ readers