12 best LLM serving GitHub repos to use
Run, serve, and route models yourself. Local inference, gateways, and the proxies that let you swap providers without rewriting your app.
Find GitHub repos worth using by topic, language, license and how active they are. Navid's picks come first, and each one opens its own page.
This is a list of the best LLM serving GitHub repos.
In fact, it has 12 of them, with Navid's picks first.
So if you want LLM serving GitHub repos worth your time, you'll love this list.
Run, serve, and route models yourself. Local inference, gateways, and the proxies that let you swap providers without rewriting your app.
Here's what's inside:
- Ollama by Ollama
- llama.cpp by ggml-org
- vLLM by vllm-project
- Unsloth by unslothai
- GPT4Free by xtekky
- LiteLLM by BerriAI
- AirLLM by lyogavin
- Gitleaks by gitleaks
- Dyad by dyad-sh
- Skypilot by skypilot-org
- ClawRouter by BlockRunAI
- LM Studio CLI by LM Studio
Each one comes with what it covers and who it's for.
What are the best LLM serving GitHub repos?
Here's the list at a glance.
- Owner
- Ollama
- Stars
- ★ 181k
- Owner
- ggml-org
- Stars
- ★ 129k
- Owner
- vllm-project
- Stars
- ★ 93k
- Owner
- unslothai
- Stars
- ★ 77k
- Owner
- xtekky
- Stars
- ★ 67k
- Owner
- BerriAI
- Stars
- ★ 59k
- Owner
- lyogavin
- Stars
- ★ 35k
- Owner
- gitleaks
- Stars
- ★ 29k
- Owner
- dyad-sh
- Stars
- ★ 22k
- Owner
- skypilot-org
- Stars
- ★ 11k
- Owner
- BlockRunAI
- Stars
- ★ 6.6k
- Owner
- LM Studio
- Stars
- ★ 5.3k
Top 12 LLM serving GitHub repos
1. Ollama by Ollama
The official Ollama Docker image ollama/ollama is available on Docker Hub. ollama-python ollama-js Discord 𝕏 (Twitter) Reddit
You'll be prompted to run a model or connect Ollama to your existing agents or applications such as Claude Code, OpenClaw, OpenCode, Codex, Copilot, and more.
Supported integrations include Claude Code, Codex, Copilot CLI, DeepSeek Harness, Droid, and OpenCode.
Stars: 181k
Language: Go
License: MIT
Install:
bash curl -fsSL https://ollama.com/install.sh | sh
View Ollama on GitHub · More about Ollama
2. llama.cpp by ggml-org
A few options to get llama.cpp installed on your machine: Visit https://llama.app and follow the instructions Run with Docker - see our Docker documentation Download pre-built binaries from the releases page Build from source by cloning this repository - check out our build guide
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud. Plain C/C++ implementation without any dependencies Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks AVX, AVX2, AVX512 and AMX support for x86 architectures RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA) Vulkan and SYCL backend support CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
Stars: 129k
Language: C++
License: MIT
View llama.cpp on GitHub · More about llama.cpp
3. vLLM by vllm-project
We have built a vLLM website to help you get started with vLLM. Please visit vllm.ai to learn more. For events, please visit vllm.ai/events to join us.
vLLM is a fast and easy-to-use library for LLM inference and serving.
Originally developed in the Sky Computing Lab at UC Berkeley, vLLM has grown into one of the most active open-source AI projects built and maintained by a diverse community of many dozens of academic institutions and companies from over 2000 contributors.
Stars: 93k
Language: Python
License: Apache-2.0
Install:
bash uv pip install vllm
View vLLM on GitHub · More about vLLM
4. Unsloth by unslothai
Unsloth lets you run, train, and deploy AI models locally, with support for all types of models. Run and train LLMs, diffusion, embedding, audio models: Kimi K3, MiniMax-H3, Qwen3.8, Muse Glimmer, DeepSeek-V4, Gemma 4. Agents & Tools: Use local models with Claude Code, Codex, and MCP, including tool calling and code execution. Search & RAG: Use private and unlimited web search, deep research, and RAG. Image and video: Run and train image and video diffusion or multimodal models Audio: Use private and unlimited web search, deep research, and RAG. Hardware: Supports CPU, NVIDIA, AMD, Intel, macOS, and multi GPU setups. Remote Access: Access your local models remotely through secure Cloudflare HTTPS. Fine-tuning: Train LLMs, diffusion, TTS, and embedding models 2× faster with 70% less VRAM Complete support: Supports reinforcement learning, LoRA, QLoRA, full fine tuning, pretraining, RL, GRPO, DPO, and FP8. Export & Deploy: Export or Deploy models with including GGUF, NVFP4, FP8 and more formats. Datasets: Build datasets from PDFs, CSVs, DOCX files, and more with Data Recipes. OpenAI Compatible API: Serve models through an OpenAI compatible API and also connect to cloud providers
Unsloth Start connects Claude Code, Codex and other agents to local models with one command.
Unsloth can be used in three ways: Unsloth Desktop, the desktop app; Unsloth Studio, the web UI; or Unsloth Core, the code based version.
Stars: 77k
Language: Python
License: Apache-2.0
Install:
bash curl -fsSL https://unsloth.ai/install.sh | sh
View Unsloth on GitHub · More about Unsloth
5. GPT4Free by xtekky
Live demo & docs: https://g4f.dev Documentation: https://g4f.dev/docs
GPT4Free (g4f) is a community-driven project that aggregates multiple accessible providers and interfaces to make working with modern LLMs and media-generation models easier and more flexible. GPT4Free aims to offer multi-provider support, local GUI, OpenAI-compatible REST APIs, and convenient Python and JavaScript clients, all under a community-first license.
This README is a consolidated, improved, and complete guide to installing, running, and contributing to GPT4Free.
Stars: 67k
Language: Python
License: GPL-3.0
Install:
bash docker pull hlohaus789/g4f
View GPT4Free on GitHub · More about GPT4Free
6. LiteLLM by BerriAI
Open Source AI Gateway for 100+ LLMs. Self-hosted. Enterprise-ready. Call any LLM in OpenAI format.
LiteLLM is an open source AI Gateway that gives you a single, unified interface to call 100+ LLM providers, OpenAI, Anthropic, Gemini, Bedrock, Azure, and more, using the OpenAI format.
Use it as a Python SDK for direct library integration, or deploy the AI Gateway (Proxy Server) as a centralized service for your team or organization.
Stars: 59k
Language: Python
License: Other
Install:
bash git clone https://github.com/BerriAI/litellm.git
View LiteLLM on GitHub · More about LiteLLM
7. AirLLM by lyogavin
AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card, without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB, DeepSeek-V3 (671B) on ~12GB, and Kimi K3 (2.8T), the largest open-source model released to date, on under 4GB, because sparse MoE models stream one expert at a time rather than a whole layer.
[2026/07] Kimi K3 (2.8T) support: the largest open-source model runs on a single card in 3.72GB of VRAM, measured end to end on one RTX 6000 Ada. Per-expert streaming loads only the experts a token actually routes to. K3 brings three requirements of its own: pip install compressed-tensors flash-attn (its model code mandates flash attention regardless of what you request), a CUDA 12 build of torch, since no prebuilt flash-attn wheel exists for CUDA 13 yet, and transformers 4.56.x, as its remote code does not load on 5.x.
[2026/06] v3.0: FP8 model support + the latest models. Run DeepSeek-V3 (671B) on ~12GB and Qwen3-235B on ~3GB, plus Qwen3, Llama 3.x/4, DeepSeek V2/V3, Phi-4, Gemma and more, all through a single AutoModel.
Stars: 35k
Language: Jupyter Notebook
License: Apache-2.0
Install:
bash pip install airllm
View AirLLM on GitHub · More about AirLLM
8. Gitleaks by gitleaks
[!WARNING] Gitleaks is feature complete. I'm not merging new features into Gitleaks. Future releases will be security patches only. I'm shifting my focus to Betterleaks
[![GitHub Action Test][badge-build]][build] [![Docker Hub][dockerhub-badge]][dockerhub] [![Gitleaks Action][gitleaks-badge]][gitleaks-action] [![GoDoc][go-docs-badge]][go-docs] [![GoReportCard][go-report-card-badge]][go-report-card] [![License][badge-license]][license]
Gitleaks is a tool for detecting secrets like passwords, API keys, and tokens in git repos, files, and whatever else you wanna throw at it via stdin. If you wanna learn more about how the detection engine works check out this blog: Regex is (almost) all you need.
Stars: 29k
Language: Go
License: MIT
Install:
bash brew install gitleaks
View Gitleaks on GitHub · More about Gitleaks
9. Dyad by dyad-sh
Dyad is a local, open-source AI app builder. It's fast, private, and fully under your control, like Lovable, v0, or Bolt, but running right on your machine.
More info at: https://dyad.sh/ Local: Fast, private and no lock-in. Bring your own keys: Use your own AI API keys, no vendor lock-in. Cross-platform: Easy to run on Mac or Windows.
Join our growing community of AI app builders on Reddit: r/dyadbuilders - share your projects and get help from the community!
Stars: 22k
Language: TypeScript
License: Other
View Dyad on GitHub · More about Dyad
10. Skypilot by skypilot-org
SkyPilot is a system to run, manage, and scale AI workloads on any AI infrastructure.
SkyPilot gives AI teams a simple interface to run jobs on any infra. Infra teams get a unified control plane to manage any AI compute, with advanced scheduling, scaling, and orchestration.
fire: News:fire: [Aug 2026] RL is bottlenecked by inference: scale it independently with SkyPilot: blog [Jul 2026] Serving Kimi K3 on your own GPUs with SkyPilot: blog [Jul 2026] SkyPilot v0.13.0 released: Hugging Face storage, batch inference abstractions, lifecycle hooks, governance & robustness on API server: Release notes [Jun 2026] SkyPilot Endpoints: production-ready inference on every cluster you own: blog [Jun 2026] Announcing SkyPilot Sandboxes: run untrusted, LLM-generated code on the Kubernetes clusters you already own. Learn more, join early access [May 2026] How Multiverse doubled their GPU utilization with SkyPilot: case study [Apr 2026] Introducing GPU Compass: One dashboard to browse, compare pricing, and launch across every GPU cloud. Try it at gpus.skypilot.co. [Apr 2026] Research-Driven Agents: Agents read arxiv papers before coding, landed 5 llama.cpp kernel fusions and +15% faster flash attention in ~3 hours for ~$29: blog, HackerNews [Mar 2026] Scaling Karpathy's Autoresearch: Autoresearch runs 1 experiment at a time. We gave it 16 GPUs and let it run in parallel: blog, HackerNews [Mar 2026] How H Company Unlocked Online RL and Unified their AI Platform: case study
Stars: 11k
Language: Python
License: Apache-2.0
Install:
bash pip install -r requirements.txt
View Skypilot on GitHub · More about Skypilot
11. ClawRouter by BlockRunAI
Agents can't sign up for accounts. Agents can't enter credit cards. Agents can only sign transactions. ClawRouter is the only LLM router that lets agents operate independently. 5 models free, no crypto required. No signup. No API key. No credit card.
ClawRouter is an open-source smart LLM router that reduces AI API costs by up to 88%. It analyzes each request across 15 dimensions and routes to the cheapest capable model in under 1ms, entirely locally. ClawRouter is the only LLM router built for autonomous AI agents, it uses wallet signatures for authentication (no API keys) and USDC micropayments via the x402 protocol (no credit cards). 70 models from OpenAI, Anthropic, Google, xAI, DeepSeek, and more. MIT licensed.
Every other LLM router was built for human developers, create an account, get an API key, pick a model from a dashboard, pay with a credit card.
Stars: 6.6k
Language: TypeScript
License: MIT
Install:
bash npm install -g @blockrun/clawrouter
View ClawRouter on GitHub · More about ClawRouter
12. LM Studio CLI by LM Studio
lms - Command Line Tool for LM Studio Built with lmstudio.js
If you have trouble running the command, try running npx lmstudio install-cli to add it to path.
You can use lms --help to see a list of all available subcommands.
Stars: 5.3k
Language: TypeScript
License: MIT
Install:
bash git clone https://github.com/lmstudio-ai/lmstudio-js.git --recursive
Final thoughts on LLM serving GitHub repos
The best LLM serving repo is the one you come back to. Pick 2 or 3 from this list, give them a month, and keep the ones you look forward to.
Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.
More GitHub repos
More repo topics
More free tools
Related MCP servers & CLIs
The most actionable AI newsletter for founders
Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.
No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.
P.S. Sign up now to get free access to my ultimate AI tools guide for creators.



















































