Navid MoazzezNavid Moazzez

12 best LLM serving GitHub repos to use

Run, serve, and route models yourself. Local inference, gateways, and the proxies that let you swap providers without rewriting your app.

Navid Moazzezby Navid Moazzez·Updated Sept 30, 2026·8 min read

Find GitHub repos worth using by topic, language, license and how active they are. Navid's picks come first, and each one opens its own page.

This is a list of the best LLM serving GitHub repos.

In fact, it has 12 of them, with Navid's picks first.

So if you want LLM serving GitHub repos worth your time, you'll love this list.

Run, serve, and route models yourself. Local inference, gateways, and the proxies that let you swap providers without rewriting your app.

Here's what's inside:

Each one comes with what it covers and who it's for.

What are the best LLM serving GitHub repos?

Here's the list at a glance.

Owner
Ollama
Stars
★ 181k
Owner
ggml-org
Stars
★ 129k
Owner
vllm-project
Stars
★ 93k
Owner
unslothai
Stars
★ 77k
Owner
xtekky
Stars
★ 67k
Owner
BerriAI
Stars
★ 59k
Owner
lyogavin
Stars
★ 35k
Owner
gitleaks
Stars
★ 29k
Owner
dyad-sh
Stars
★ 22k
Owner
skypilot-org
Stars
★ 11k
Owner
BlockRunAI
Stars
★ 6.6k
Owner
LM Studio
Stars
★ 5.3k

Top 12 LLM serving GitHub repos

1. Ollama by Ollama

The official Ollama Docker image ollama/ollama is available on Docker Hub. ollama-python ollama-js Discord 𝕏 (Twitter) Reddit

You'll be prompted to run a model or connect Ollama to your existing agents or applications such as Claude Code, OpenClaw, OpenCode, Codex, Copilot, and more.

Supported integrations include Claude Code, Codex, Copilot CLI, DeepSeek Harness, Droid, and OpenCode.

Stars: 181k

Language: Go

License: MIT

Install:

bash curl -fsSL https://ollama.com/install.sh | sh 

View Ollama on GitHub · More about Ollama

2. llama.cpp by ggml-org

A few options to get llama.cpp installed on your machine: Visit https://llama.app and follow the instructions Run with Docker - see our Docker documentation Download pre-built binaries from the releases page Build from source by cloning this repository - check out our build guide

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud. Plain C/C++ implementation without any dependencies Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks AVX, AVX2, AVX512 and AMX support for x86 architectures RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA) Vulkan and SYCL backend support CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

Stars: 129k

Language: C++

License: MIT

View llama.cpp on GitHub · More about llama.cpp

3. vLLM by vllm-project

We have built a vLLM website to help you get started with vLLM. Please visit vllm.ai to learn more. For events, please visit vllm.ai/events to join us.

vLLM is a fast and easy-to-use library for LLM inference and serving.

Originally developed in the Sky Computing Lab at UC Berkeley, vLLM has grown into one of the most active open-source AI projects built and maintained by a diverse community of many dozens of academic institutions and companies from over 2000 contributors.

Stars: 93k

Language: Python

License: Apache-2.0

Install:

bash uv pip install vllm 

View vLLM on GitHub · More about vLLM

4. Unsloth by unslothai

Unsloth lets you run, train, and deploy AI models locally, with support for all types of models. Run and train LLMs, diffusion, embedding, audio models: Kimi K3, MiniMax-H3, Qwen3.8, Muse Glimmer, DeepSeek-V4, Gemma 4. Agents & Tools: Use local models with Claude Code, Codex, and MCP, including tool calling and code execution. Search & RAG: Use private and unlimited web search, deep research, and RAG. Image and video: Run and train image and video diffusion or multimodal models Audio: Use private and unlimited web search, deep research, and RAG. Hardware: Supports CPU, NVIDIA, AMD, Intel, macOS, and multi GPU setups. Remote Access: Access your local models remotely through secure Cloudflare HTTPS. Fine-tuning: Train LLMs, diffusion, TTS, and embedding models 2× faster with 70% less VRAM Complete support: Supports reinforcement learning, LoRA, QLoRA, full fine tuning, pretraining, RL, GRPO, DPO, and FP8. Export & Deploy: Export or Deploy models with including GGUF, NVFP4, FP8 and more formats. Datasets: Build datasets from PDFs, CSVs, DOCX files, and more with Data Recipes. OpenAI Compatible API: Serve models through an OpenAI compatible API and also connect to cloud providers

Unsloth Start connects Claude Code, Codex and other agents to local models with one command.

Unsloth can be used in three ways: Unsloth Desktop, the desktop app; Unsloth Studio, the web UI; or Unsloth Core, the code based version.

Stars: 77k

Language: Python

License: Apache-2.0

Install:

bash curl -fsSL https://unsloth.ai/install.sh | sh 

View Unsloth on GitHub · More about Unsloth

5. GPT4Free by xtekky

Live demo & docs: https://g4f.dev Documentation: https://g4f.dev/docs

GPT4Free (g4f) is a community-driven project that aggregates multiple accessible providers and interfaces to make working with modern LLMs and media-generation models easier and more flexible. GPT4Free aims to offer multi-provider support, local GUI, OpenAI-compatible REST APIs, and convenient Python and JavaScript clients, all under a community-first license.

This README is a consolidated, improved, and complete guide to installing, running, and contributing to GPT4Free.

Stars: 67k

Language: Python

License: GPL-3.0

Install:

bash docker pull hlohaus789/g4f 

View GPT4Free on GitHub · More about GPT4Free

6. LiteLLM by BerriAI

Open Source AI Gateway for 100+ LLMs. Self-hosted. Enterprise-ready. Call any LLM in OpenAI format.

LiteLLM is an open source AI Gateway that gives you a single, unified interface to call 100+ LLM providers, OpenAI, Anthropic, Gemini, Bedrock, Azure, and more, using the OpenAI format.

Use it as a Python SDK for direct library integration, or deploy the AI Gateway (Proxy Server) as a centralized service for your team or organization.

Stars: 59k

Language: Python

License: Other

Install:

bash git clone https://github.com/BerriAI/litellm.git

View LiteLLM on GitHub · More about LiteLLM

7. AirLLM by lyogavin

AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card, without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB, DeepSeek-V3 (671B) on ~12GB, and Kimi K3 (2.8T), the largest open-source model released to date, on under 4GB, because sparse MoE models stream one expert at a time rather than a whole layer.

[2026/07] Kimi K3 (2.8T) support: the largest open-source model runs on a single card in 3.72GB of VRAM, measured end to end on one RTX 6000 Ada. Per-expert streaming loads only the experts a token actually routes to. K3 brings three requirements of its own: pip install compressed-tensors flash-attn (its model code mandates flash attention regardless of what you request), a CUDA 12 build of torch, since no prebuilt flash-attn wheel exists for CUDA 13 yet, and transformers 4.56.x, as its remote code does not load on 5.x.

[2026/06] v3.0: FP8 model support + the latest models. Run DeepSeek-V3 (671B) on ~12GB and Qwen3-235B on ~3GB, plus Qwen3, Llama 3.x/4, DeepSeek V2/V3, Phi-4, Gemma and more, all through a single AutoModel.

Stars: 35k

Language: Jupyter Notebook

License: Apache-2.0

Install:

bash pip install airllm 

View AirLLM on GitHub · More about AirLLM

8. Gitleaks by gitleaks

[!WARNING] Gitleaks is feature complete. I'm not merging new features into Gitleaks. Future releases will be security patches only. I'm shifting my focus to Betterleaks

[![GitHub Action Test][badge-build]][build] [![Docker Hub][dockerhub-badge]][dockerhub] [![Gitleaks Action][gitleaks-badge]][gitleaks-action] [![GoDoc][go-docs-badge]][go-docs] [![GoReportCard][go-report-card-badge]][go-report-card] [![License][badge-license]][license]

Gitleaks is a tool for detecting secrets like passwords, API keys, and tokens in git repos, files, and whatever else you wanna throw at it via stdin. If you wanna learn more about how the detection engine works check out this blog: Regex is (almost) all you need.

Stars: 29k

Language: Go

License: MIT

Install:

bash brew install gitleaks 

View Gitleaks on GitHub · More about Gitleaks

9. Dyad by dyad-sh

Dyad is a local, open-source AI app builder. It's fast, private, and fully under your control, like Lovable, v0, or Bolt, but running right on your machine.

More info at: https://dyad.sh/ Local: Fast, private and no lock-in. Bring your own keys: Use your own AI API keys, no vendor lock-in. Cross-platform: Easy to run on Mac or Windows.

Join our growing community of AI app builders on Reddit: r/dyadbuilders - share your projects and get help from the community!

Stars: 22k

Language: TypeScript

License: Other

View Dyad on GitHub · More about Dyad

10. Skypilot by skypilot-org

SkyPilot is a system to run, manage, and scale AI workloads on any AI infrastructure.

SkyPilot gives AI teams a simple interface to run jobs on any infra. Infra teams get a unified control plane to manage any AI compute, with advanced scheduling, scaling, and orchestration.

fire: News:fire: [Aug 2026] RL is bottlenecked by inference: scale it independently with SkyPilot: blog [Jul 2026] Serving Kimi K3 on your own GPUs with SkyPilot: blog [Jul 2026] SkyPilot v0.13.0 released: Hugging Face storage, batch inference abstractions, lifecycle hooks, governance & robustness on API server: Release notes [Jun 2026] SkyPilot Endpoints: production-ready inference on every cluster you own: blog [Jun 2026] Announcing SkyPilot Sandboxes: run untrusted, LLM-generated code on the Kubernetes clusters you already own. Learn more, join early access [May 2026] How Multiverse doubled their GPU utilization with SkyPilot: case study [Apr 2026] Introducing GPU Compass: One dashboard to browse, compare pricing, and launch across every GPU cloud. Try it at gpus.skypilot.co. [Apr 2026] Research-Driven Agents: Agents read arxiv papers before coding, landed 5 llama.cpp kernel fusions and +15% faster flash attention in ~3 hours for ~$29: blog, HackerNews [Mar 2026] Scaling Karpathy's Autoresearch: Autoresearch runs 1 experiment at a time. We gave it 16 GPUs and let it run in parallel: blog, HackerNews [Mar 2026] How H Company Unlocked Online RL and Unified their AI Platform: case study

Stars: 11k

Language: Python

License: Apache-2.0

Install:

bash pip install -r requirements.txt 

View Skypilot on GitHub · More about Skypilot

11. ClawRouter by BlockRunAI

Agents can't sign up for accounts. Agents can't enter credit cards. Agents can only sign transactions. ClawRouter is the only LLM router that lets agents operate independently. 5 models free, no crypto required. No signup. No API key. No credit card.

ClawRouter is an open-source smart LLM router that reduces AI API costs by up to 88%. It analyzes each request across 15 dimensions and routes to the cheapest capable model in under 1ms, entirely locally. ClawRouter is the only LLM router built for autonomous AI agents, it uses wallet signatures for authentication (no API keys) and USDC micropayments via the x402 protocol (no credit cards). 70 models from OpenAI, Anthropic, Google, xAI, DeepSeek, and more. MIT licensed.

Every other LLM router was built for human developers, create an account, get an API key, pick a model from a dashboard, pay with a credit card.

Stars: 6.6k

Language: TypeScript

License: MIT

Install:

bash npm install -g @blockrun/clawrouter 

View ClawRouter on GitHub · More about ClawRouter

12. LM Studio CLI by LM Studio

lms - Command Line Tool for LM Studio Built with lmstudio.js

If you have trouble running the command, try running npx lmstudio install-cli to add it to path.

You can use lms --help to see a list of all available subcommands.

Stars: 5.3k

Language: TypeScript

License: MIT

Install:

bash git clone https://github.com/lmstudio-ai/lmstudio-js.git --recursive 

View LM Studio CLI on GitHub · More about LM Studio CLI

Final thoughts on LLM serving GitHub repos

The best LLM serving repo is the one you come back to. Pick 2 or 3 from this list, give them a month, and keep the ones you look forward to.

Navid Moazzez

AI business strategist & AI OS builder

Navid Moazzez helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life.

Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.

More GitHub repos

More free tools

Related MCP servers & CLIs

Free AI newsletter

The most actionable AI newsletter for founders

Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.

No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.

P.S. Sign up now to get free access to my ultimate AI tools guide for creators.

Loved by 10,000+ readers