llama.cpp
A few options to get llama.cpp installed on your machine: Visit https://llama.app and follow the instructions Run with Docker - see our Docker documentation Download pre-built binaries from the releases page Build from source by cloning this repository - check out our build guide The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud. Plain C/C++ implementation without any dependencies Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks AVX, AVX2, AVX512 and AMX support for x86 architectures RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA) Vulkan and SYCL backend support CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
View llama.cpp on GitHub
A few options to get llama.cpp installed on your machine: Visit https://llama.app and follow the instructions Run with Docker - see our Docker documentation Download pre-built binaries from the releases page Build from source by cloning this repository - check out our build guide
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud. Plain C/C++ implementation without any dependencies Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks AVX, AVX2, AVX512 and AMX support for x86 architectures RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA) Vulkan and SYCL backend support CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
llama.cpp at a glance
| Stars | 129k |
|---|---|
| Forks | 23k |
| Language | C++ |
| License | MIT |
| Last update | 2026-09-20 |
| Contributors | 2020 |
Where llama.cpp is listed
Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.
GitHub repos like this
More repo topics
More free tools
Related MCP servers & CLIs
The most actionable AI newsletter for founders
Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.
No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.
P.S. Sign up now to get free access to my ultimate AI tools guide for creators.






































