evolve

GitHub

基于达尔文模式的自进化测试框架,通过沙箱化变异、评分和选择机制自动优化测试工具链,提升模型修复能力并降低推理成本。

kimi-k3-harness/.claude/skills/evolve/SKILL.md ruvnet/metaharness

Trigger Scenarios

需要自动化优化测试执行流程 希望利用进化算法提升代码修复率 需要低成本高吞吐的测试迭代策略

Install

npx skills add ruvnet/metaharness --skill evolve -g -y
More Options

Non-standard path

npx skills add https://github.com/ruvnet/metaharness/tree/main/kimi-k3-harness/.claude/skills/evolve -g -y

Use without installing

npx skills use ruvnet/metaharness@evolve

指定 Agent (Claude Code)

npx skills add ruvnet/metaharness --skill evolve -a claude-code -g -y

安装 repo 全部 skill

npx skills add ruvnet/metaharness --all -g -y

预览 repo 内 skill

npx skills add ruvnet/metaharness --list

SKILL.md

Frontmatter
{
    "name": "evolve",
    "description": "Evolve this harness with Darwin Mode — frozen model, evolving harness (real, sandboxed, safety-gated)."
}

evolve — Darwin Mode self-improvement

kimi-k3-harness ships with Darwin Mode (@metaharness/darwin, ADR-070…146): the model is frozen; the harness evolves. Each generation mutates ONE of the 7 surface files (planner, contextBuilder, reviewer, retry/tool/memory/score policy), sandboxes each child, scores it, and keeps only variants that measurably improve — building an archive of successful descendants.

Run it

npm run evolve        # real substrate: runs your test command per variant (deterministic mutator — no API key, no network)
npm run evolve:dry    # mock substrate: fast, fully offline, no test execution

Or directly:

npx metaharness-darwin evolve . --sandbox real --generations 3 --children 4

Safety (secure by default)

  • Deterministic mutator is the default — no network, no API key, air-gapped.
  • Every mutation passes the validateGeneratedCode gate: no new imports, network, filesystem, shell, env access, or dependencies — pure refactor/tuning only.
  • Mutations run in a sandbox; only variants that pass your tests are archived.
  • Nothing is promoted without measured improvement (guard against Goodharting).

See @metaharness/darwin for selection strategies (--selection, --crossover, --curriculum), statistical gates (--fdr, --bench), and the real-LLM mutator (library API).

What the benchmarks taught us (measured, full SWE-bench Lite 300)

Defaults worth carrying into how you evolve and run this harness (full evidence + CIs in @metaharness/darwin's LEARNINGS.md / bench/results/RESULTS.md):

  1. Closed-loop repair is the #1 lever (~2×). Feeding test/compiler failure back and retrying took resolve-rate 7.7% → 15.3% on the same cheap model. Iterate against ground truth, don't single-shot.
  2. Cheap-first + cost-aware routing. Track $/resolve, not just resolve-rate; a cheap model resolved 31× cheaper per fix than a frontier one. Reserve frontier for measured capability gaps.
  3. Tier the models (Barbarian & Scholar). Cheap sweep + frontier on only the residual = 33.3% at ~6× lower cost than running frontier everywhere.
  4. Put the output-format contract in a system message + example, and size prompts to the model's real context window — this alone took a weak local model from 0% to ~50% valid output.
  5. Only trust batch evaluation of the final artifact — in-loop counters drift 1.5–5×.
  6. The harness multiplies the model; it can't rescue one below the task's reasoning floor. Pick the smallest model above the floor, then let evolution do the rest.

Version History

  • 6840275 Current 2026-08-12 13:13

Same Skill Collection

.claude-plugin/skills/compare-harnesses/SKILL.md
.claude-plugin/skills/create-harness/SKILL.md
.claude-plugin/skills/diag-harness/SKILL.md
.claude-plugin/skills/example-harness/SKILL.md
.claude-plugin/skills/list-templates/SKILL.md
.claude-plugin/skills/oia-manifest/SKILL.md
.claude-plugin/skills/publish-harness/SKILL.md
.claude-plugin/skills/repo-genome/SKILL.md
.claude-plugin/skills/upgrade-harness/SKILL.md
.claude-plugin/skills/validate-harness/SKILL.md
.claude-plugin/skills/verify-witness/SKILL.md
kimi-k3-harness/.claude/skills/plan-change/SKILL.md

Metadata

Files
0
Version
6840275
Hash
14b28a28
Indexed
2026-08-12 13:13

- 위키
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-08-12 18:18
浙ICP备14020137号-1 $방문자$