Agent Skillsruvnet/RuVector › evolve

evolve

GitHub

基于达尔文进化模式的自动化代码优化技能。通过沙箱化变异、测试验证和评分,自动迭代改进代码逻辑与配置,提升解决率并降低成本,同时确保安全性与无网络依赖。

harnesses/timesfm-harness/.claude/skills/evolve/SKILL.md ruvnet/RuVector

Trigger Scenarios

需要自动化优化代码或配置以提升性能时 执行 npm run evolve 命令进行自改进时

Install

npx skills add ruvnet/RuVector --skill evolve -g -y
More Options

Non-standard path

npx skills add https://github.com/ruvnet/RuVector/tree/main/harnesses/timesfm-harness/.claude/skills/evolve -g -y

Use without installing

npx skills use ruvnet/RuVector@evolve

指定 Agent (Claude Code)

npx skills add ruvnet/RuVector --skill evolve -a claude-code -g -y

安装 repo 全部 skill

npx skills add ruvnet/RuVector --all -g -y

预览 repo 内 skill

npx skills add ruvnet/RuVector --list

SKILL.md

Frontmatter
{
    "name": "evolve",
    "description": "Evolve this harness with Darwin Mode — frozen model, evolving harness (real, sandboxed, safety-gated)."
}

evolve — Darwin Mode self-improvement

timesfm-harness ships with Darwin Mode (@metaharness/darwin, ADR-070…146): the model is frozen; the harness evolves. Each generation mutates ONE of the 7 surface files (planner, contextBuilder, reviewer, retry/tool/memory/score policy), sandboxes each child, scores it, and keeps only variants that measurably improve — building an archive of successful descendants.

Run it

npm run evolve        # real substrate: runs your test command per variant (deterministic mutator — no API key, no network)
npm run evolve:dry    # mock substrate: fast, fully offline, no test execution

Or directly:

npx metaharness-darwin evolve . --sandbox real --generations 3 --children 4

Safety (secure by default)

  • Deterministic mutator is the default — no network, no API key, air-gapped.
  • Every mutation passes the validateGeneratedCode gate: no new imports, network, filesystem, shell, env access, or dependencies — pure refactor/tuning only.
  • Mutations run in a sandbox; only variants that pass your tests are archived.
  • Nothing is promoted without measured improvement (guard against Goodharting).

See @metaharness/darwin for selection strategies (--selection, --crossover, --curriculum), statistical gates (--fdr, --bench), and the real-LLM mutator (library API).

What the benchmarks taught us (measured, full SWE-bench Lite 300)

Defaults worth carrying into how you evolve and run this harness (full evidence + CIs in @metaharness/darwin's LEARNINGS.md / bench/results/RESULTS.md):

  1. Closed-loop repair is the #1 lever (~2×). Feeding test/compiler failure back and retrying took resolve-rate 7.7% → 15.3% on the same cheap model. Iterate against ground truth, don't single-shot.
  2. Cheap-first + cost-aware routing. Track $/resolve, not just resolve-rate; a cheap model resolved 31× cheaper per fix than a frontier one. Reserve frontier for measured capability gaps.
  3. Tier the models (Barbarian & Scholar). Cheap sweep + frontier on only the residual = 33.3% at ~6× lower cost than running frontier everywhere.
  4. Put the output-format contract in a system message + example, and size prompts to the model's real context window — this alone took a weak local model from 0% to ~50% valid output.
  5. Only trust batch evaluation of the final artifact — in-loop counters drift 1.5–5×.
  6. The harness multiplies the model; it can't rescue one below the task's reasoning floor. Pick the smallest model above the floor, then let evolution do the rest.

Version History

  • 74d2a60 Current 2026-08-20 07:51

Same Skill Collection

.claude/skills/agentdb-advanced/SKILL.md
.claude/skills/agentdb-learning/SKILL.md
.claude/skills/agentdb-memory-patterns/SKILL.md
.claude/skills/agentdb-optimization/SKILL.md
.claude/skills/agentdb-vector-search/SKILL.md
.claude/skills/agentic-jujutsu/SKILL.md
.claude/skills/browser/SKILL.md
.claude/skills/custom-workers/SKILL.md
.claude/skills/flow-nexus-neural/SKILL.md
.claude/skills/flow-nexus-platform/SKILL.md
.claude/skills/flow-nexus-swarm/SKILL.md
.claude/skills/github-code-review/SKILL.md
.claude/skills/github-multi-repo/SKILL.md
.claude/skills/github-project-management/SKILL.md
.claude/skills/github-release-management/SKILL.md
.claude/skills/github-workflow-automation/SKILL.md
.claude/skills/hive-mind-advanced/SKILL.md
.claude/skills/hooks-automation/SKILL.md
.claude/skills/pair-programming/SKILL.md
.claude/skills/performance-analysis/SKILL.md
.claude/skills/reasoningbank-agentdb/SKILL.md
.claude/skills/reasoningbank-intelligence/SKILL.md
.claude/skills/skill-builder/SKILL.md
.claude/skills/sparc-methodology/SKILL.md
.claude/skills/stream-chain/SKILL.md
.claude/skills/swarm-advanced/SKILL.md
.claude/skills/swarm-orchestration/SKILL.md
.claude/skills/v3-cli-modernization/SKILL.md
.claude/skills/v3-core-implementation/SKILL.md
.claude/skills/v3-ddd-architecture/SKILL.md
.claude/skills/v3-integration-deep/SKILL.md
.claude/skills/v3-mcp-optimization/SKILL.md
.claude/skills/v3-memory-unification/SKILL.md
.claude/skills/v3-performance-optimization/SKILL.md
.claude/skills/v3-security-overhaul/SKILL.md
.claude/skills/v3-swarm-coordination/SKILL.md
.claude/skills/verification-quality/SKILL.md
harnesses/timesfm-harness/.claude/skills/plan-change/SKILL.md
ui/ruvocal/.claude/skills/add-model-descriptions/SKILL.md

Metadata

Files
0
Version
74d2a60
Hash
298f66d6
Indexed
2026-08-20 07:51

- 위키
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-08-26 10:46
浙ICP备14020137号-1 $방문자$