Skill Up
skill-up is an evaluation and evolution tool for Agent Skills. Evaluation makes Skill quality measurable and repeatable: declarative YAML cases run across multiple Agent Engines, use rule, script, or Agent judges, and produce structured reports locally or in CI. Evolution turns those results into the next improvement: through conversation, skill-upper reads failures, automatically repairs or expands the eval suite, reruns skill-up, and keeps iterating with you. Eval-to-Evolution Loop with skill-upper: Create evals through natural conversation, diagnose failures, automatically repair or expand cases, and rerun skill-up until the eval suite evolves. Declarative Eval Config: Define evaluation environment, engine, model, and cases through YAML (eval.yaml + cases/.yaml). Multi-Engine Support: Works with Qoder CLI, Claude Code, and Codex as built-in Agent Engines, plus user-defined agents via engine.custom (local transport, see docs/design/custom-engine.md). Flexible Judging: Supports rulebased, script, and agentjudge evaluation strategies. Structured Reports: Outputs Anthropic-compatible grading.json, benchmark.json, benchmark.md, plus result.json, JUnit XML, and HTML reports. Anthropic Compatible: Import evals.json via skill-up import, or auto-detect with --auto. CI-Ready: Designed for local development and continuous integration pipelines. The official Agent Skills evaluation guide describes the right evaluation loop: write realistic cases, run with and without the Skill, grade outputs, aggregate results, and iterate. skill-up turns that workflow into a reusable CLI: Replaces ad hoc run folders with a declarative eval.yaml + cases/.yaml format. Closes the improvement loop: skill-upper can interpret failed reports, repair or add eval cases, and drive the next skill-up run through conversation. Automates workspace setup, Skill installation, Agent Engine invocation, judging, and report generation. Supports multiple engines (claudecode, codex, qodercli, qwencode) instead of tying the workflow to one client. Keeps compatibility with Anthropic-style evals.json while adding richer judges, CI-friendly commands, and structured reports. The recommended way to use skill-up is through skill-upper, the Agent Skill shipped in this repository. It lets your AI agent create evals, run skill-up, understand failures, fix the Skill or its evals, add regression coverage, and repeat the loop through conversation.
View Skill Up on GitHub
skill-up is an evaluation and evolution tool for Agent Skills. Evaluation makes Skill quality measurable and repeatable: declarative YAML cases run across multiple Agent Engines, use rule, script, or Agent judges, and produce structured reports locally or in CI. Evolution turns those results into the next improvement: through conversation, skill-upper reads failures, automatically repairs or expands the eval suite, reruns skill-up, and keeps iterating with you. Eval-to-Evolution Loop with skill-upper: Create evals through natural conversation, diagnose failures, automatically repair or expand cases, and rerun skill-up until the eval suite evolves. Declarative Eval Config: Define evaluation environment, engine, model, and cases through YAML (eval.yaml + cases/.yaml). Multi-Engine Support: Works with Qoder CLI, Claude Code, and Codex as built-in Agent Engines, plus user-defined agents via engine.custom (local transport, see docs/design/custom-engine.md). Flexible Judging: Supports rulebased, script, and agentjudge evaluation strategies. Structured Reports: Outputs Anthropic-compatible grading.json, benchmark.json, benchmark.md, plus result.json, JUnit XML, and HTML reports. Anthropic Compatible: Import evals.json via skill-up import, or auto-detect with --auto. CI-Ready: Designed for local development and continuous integration pipelines.
The official Agent Skills evaluation guide describes the right evaluation loop: write realistic cases, run with and without the Skill, grade outputs, aggregate results, and iterate. skill-up turns that workflow into a reusable CLI: Replaces ad hoc run folders with a declarative eval.yaml + cases/.yaml format. Closes the improvement loop: skill-upper can interpret failed reports, repair or add eval cases, and drive the next skill-up run through conversation. Automates workspace setup, Skill installation, Agent Engine invocation, judging, and report generation. Supports multiple engines (claudecode, codex, qodercli, qwencode) instead of tying the workflow to one client. Keeps compatibility with Anthropic-style evals.json while adding richer judges, CI-friendly commands, and structured reports.
The recommended way to use skill-up is through skill-upper, the Agent Skill shipped in this repository. It lets your AI agent create evals, run skill-up, understand failures, fix the Skill or its evals, add regression coverage, and repeat the loop through conversation.
Skill Up at a glance
| Stars | 1.1k |
|---|---|
| Forks | 91 |
| Language | Go |
| License | Apache-2.0 |
| Last update | 2026-09-28 |
| Contributors | 13 |
How to install Skill Up
bash curl -fsSL https://raw.githubusercontent.com/alibaba/skill-up/main/install.sh | bash
Where Skill Up is listed
Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.
GitHub repos like this
More repo topics
More free tools
Related MCP servers & CLIs
The most actionable AI newsletter for founders
Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.
No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.
P.S. Sign up now to get free access to my ultimate AI tools guide for creators.







































