Headroom
AI agents / LLMs: read /llms.txt here, or fetch the live index / full docs blob. Headroom compresses everything your AI agent reads, tool outputs, logs, RAG chunks, files, and conversation history, before it reaches the LLM. Same answers, fraction of the tokens. Live: 10,144 1,260 tokens, same FATAL found. Library, compress(messages) in Python or TypeScript, inline in any app Proxy, headroom proxy --port 8787, zero code changes, any language Agent wrap, headroom wrap claudecodexgrokcopilotcursoraideropencodeclinecontinuegooseopenhandsopenclawvibeompzcode in one command; undo with headroom unwrap MCP server, headroomcompress, headroomretrieve, headroomstats for any MCP client Cross-agent memory, shared store across Claude, Codex, Gemini, Grok, auto-dedup headroom learn, mines failed sessions, writes corrections to CLAUDE.local.md (default, gitignored) or CLAUDE.md / AGENTS.md / GEMINI.md / GROK.md Output token reduction, trims what the model writes back (not just what you send): drops ceremony/restated code and skips deep "thinking" on routine steps. See Output token reduction. Reversible (CCR), originals are cached for retrieval on demand ContentRouter, detects content type, selects the right compressor SmartCrusher / CodeCompressor / Kompress-v2-base, compress JSON, AST, or prose CacheAligner - detects and warns about volatile content that can bust provider KV cache prefixes; never rewrites prompts CCR, stores originals locally; LLM calls headroomretrieve if it needs them
View Headroom on GitHub
AI agents / LLMs: read /llms.txt here, or fetch the live index / full docs blob.
Headroom compresses everything your AI agent reads, tool outputs, logs, RAG chunks, files, and conversation history, before it reaches the LLM. Same answers, fraction of the tokens.
Live: 10,144 1,260 tokens, same FATAL found. Library, compress(messages) in Python or TypeScript, inline in any app Proxy, headroom proxy --port 8787, zero code changes, any language Agent wrap, headroom wrap claudecodexgrokcopilotcursoraideropencodeclinecontinuegooseopenhandsopenclawvibeompzcode in one command; undo with headroom unwrap MCP server, headroomcompress, headroomretrieve, headroomstats for any MCP client Cross-agent memory, shared store across Claude, Codex, Gemini, Grok, auto-dedup headroom learn, mines failed sessions, writes corrections to CLAUDE.local.md (default, gitignored) or CLAUDE.md / AGENTS.md / GEMINI.md / GROK.md Output token reduction, trims what the model writes back (not just what you send): drops ceremony/restated code and skips deep "thinking" on routine steps. See Output token reduction. Reversible (CCR), originals are cached for retrieval on demand ContentRouter, detects content type, selects the right compressor SmartCrusher / CodeCompressor / Kompress-v2-base, compress JSON, AST, or prose CacheAligner - detects and warns about volatile content that can bust provider KV cache prefixes; never rewrites prompts CCR, stores originals locally; LLM calls headroomretrieve if it needs them
Headroom at a glance
| Stars | 74k |
|---|---|
| Forks | 5.7k |
| Language | Python |
| License | Apache-2.0 |
| Last update | 2026-09-30 |
| Contributors | 294 |
How to install Headroom
bash uv tool install --python 3.13 "headroom-ai[all]" # CLI, isolated app env
Where Headroom is listed
Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.
GitHub repos like this
More repo topics
More free tools
Related MCP servers & CLIs
The most actionable AI newsletter for founders
Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.
No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.
P.S. Sign up now to get free access to my ultimate AI tools guide for creators.







































