eval-harness

GitHub

评估 Claude Code 会话的正式框架,实现评估驱动开发(EDD)。用于定义通过/失败标准、测量代理可靠性及创建回归测试套件。支持能力与回归评估,提供代码、模型及人工评分器,并采用 pass@k 等指标衡量成功率。

skills/eval-harness/SKILL.md AjayIrkal23/agentic-mercy-10x

触发场景

设置 AI 辅助工作流的评估驱动开发 定义任务完成的标准 测量代理可靠性 创建回归测试套件 跨模型版本基准测试

安装

npx skills add AjayIrkal23/agentic-mercy-10x --skill eval-harness -g -y
更多选项

不安装直接使用

npx skills use AjayIrkal23/agentic-mercy-10x@eval-harness

指定 Agent (Claude Code)

npx skills add AjayIrkal23/agentic-mercy-10x --skill eval-harness -a claude-code -g -y

安装 repo 全部 skill

npx skills add AjayIrkal23/agentic-mercy-10x --all -g -y

预览 repo 内 skill

npx skills add AjayIrkal23/agentic-mercy-10x --list

SKILL.md

Frontmatter
{
    "name": "eval-harness",
    "tools": "Read, Write, Edit, Bash, Grep, Glob",
    "origin": "ECC",
    "schema": 1,
    "category": "general",
    "surfaces": [
        "general"
    ],
    "triggers": {
        "paths": [],
        "intents": [
            "general"
        ],
        "keywords": [
            "claude",
            "code",
            "development",
            "edd",
            "eval",
            "eval-driven",
            "evaluation",
            "formal",
            "framework",
            "harness",
            "implementing",
            "principles",
            "sessions"
        ]
    },
    "platforms": [
        "linux",
        "darwin",
        "windows"
    ],
    "token-cost": 1573,
    "description": "ALWAYS invoke when evaluating Claude Code sessions — formal evaluation framework implementing eval-driven development (EDD) principles.",
    "disable-model-invocation": false
}

Eval Harness Skill

A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.

When to Activate

  • Setting up eval-driven development (EDD) for AI-assisted workflows
  • Defining pass/fail criteria for Claude Code task completion
  • Measuring agent reliability with pass@k metrics
  • Creating regression test suites for prompt or agent changes
  • Benchmarking agent performance across model versions

Philosophy

Eval-Driven Development treats evals as the "unit tests of AI development":

  • Define expected behavior BEFORE implementation
  • Run evals continuously during development
  • Track regressions with each change
  • Use pass@k metrics for reliability measurement

Eval Types

Capability Evals

Test if Claude can do something it couldn't before:

[CAPABILITY EVAL: feature-name]
Task: Description of what Claude should accomplish
Success Criteria:
  - [ ] Criterion 1
  - [ ] Criterion 2
  - [ ] Criterion 3
Expected Output: Description of expected result

Regression Evals

Ensure changes don't break existing functionality:

[REGRESSION EVAL: feature-name]
Baseline: SHA or checkpoint name
Tests:
  - existing-test-1: PASS/FAIL
  - existing-test-2: PASS/FAIL
  - existing-test-3: PASS/FAIL
Result: X/Y passed (previously Y/Y)

Grader Types

1. Code-Based Grader

Deterministic checks using code:

# Check if file contains expected pattern
grep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"

# Check if tests pass
npm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"

# Check if build succeeds
npm run build && echo "PASS" || echo "FAIL"

2. Model-Based Grader

Use Claude to evaluate open-ended outputs:

[MODEL GRADER PROMPT]
Evaluate the following code change:
1. Does it solve the stated problem?
2. Is it well-structured?
3. Are edge cases handled?
4. Is error handling appropriate?

Score: 1-5 (1=poor, 5=excellent)
Reasoning: [explanation]

3. Human Grader

Flag for manual review:

[HUMAN REVIEW REQUIRED]
Change: Description of what changed
Reason: Why human review is needed
Risk Level: LOW/MEDIUM/HIGH

Metrics

pass@k

"At least one success in k attempts"

  • pass@1: First attempt success rate
  • pass@3: Success within 3 attempts
  • Typical target: pass@3 > 90%

pass^k

"All k trials succeed"

  • Higher bar for reliability
  • pass^3: 3 consecutive successes
  • Use for critical paths

Eval Workflow

1. Define (Before Coding)

## EVAL DEFINITION: feature-xyz

### Capability Evals
1. Can create new user account
2. Can validate email format
3. Can hash password securely

### Regression Evals
1. Existing login still works
2. Session management unchanged
3. Logout flow intact

### Success Metrics
- pass@3 > 90% for capability evals
- pass^3 = 100% for regression evals

2. Implement

Write code to pass the defined evals.

3. Evaluate

# Run capability evals
[Run each capability eval, record PASS/FAIL]

# Run regression evals
npm test -- --testPathPattern="existing"

# Generate report

4. Report

EVAL REPORT: feature-xyz
========================

Capability Evals:
  create-user:     PASS (pass@1)
  validate-email:  PASS (pass@2)
  hash-password:   PASS (pass@1)
  Overall:         3/3 passed

Regression Evals:
  login-flow:      PASS
  session-mgmt:    PASS
  logout-flow:     PASS
  Overall:         3/3 passed

Metrics:
  pass@1: 67% (2/3)
  pass@3: 100% (3/3)

Status: READY FOR REVIEW

Integration Patterns

Pre-Implementation

/eval define feature-name

Creates eval definition file at .claude/evals/feature-name.md

During Implementation

/eval check feature-name

Runs current evals and reports status

Post-Implementation

/eval report feature-name

Generates full eval report

Eval Storage

Store evals in project:

.claude/
  evals/
    feature-xyz.md      # Eval definition
    feature-xyz.log     # Eval run history
    baseline.json       # Regression baselines

Best Practices

  1. Define evals BEFORE coding - Forces clear thinking about success criteria
  2. Run evals frequently - Catch regressions early
  3. Track pass@k over time - Monitor reliability trends
  4. Use code graders when possible - Deterministic > probabilistic
  5. Human review for security - Never fully automate security checks
  6. Keep evals fast - Slow evals don't get run
  7. Version evals with code - Evals are first-class artifacts

Example: Adding Authentication

## EVAL: add-authentication

### Phase 1: Define (10 min)
Capability Evals:
- [ ] User can register with email/password
- [ ] User can login with valid credentials
- [ ] Invalid credentials rejected with proper error
- [ ] Sessions persist across page reloads
- [ ] Logout clears session

Regression Evals:
- [ ] Public routes still accessible
- [ ] API responses unchanged
- [ ] Database schema compatible

### Phase 2: Implement (varies)
[Write code]

### Phase 3: Evaluate
Run: /eval check add-authentication

### Phase 4: Report
EVAL REPORT: add-authentication
==============================
Capability: 5/5 passed (pass@3: 100%)
Regression: 3/3 passed (pass^3: 100%)
Status: SHIP IT

Product Evals (v1.8)

Use product evals when behavior quality cannot be captured by unit tests alone.

Grader Types

  1. Code grader (deterministic assertions)
  2. Rule grader (regex/schema constraints)
  3. Model grader (LLM-as-judge rubric)
  4. Human grader (manual adjudication for ambiguous outputs)

pass@k Guidance

  • pass@1: direct reliability
  • pass@3: practical reliability under controlled retries
  • pass^3: stability test (all 3 runs must pass)

Recommended thresholds:

  • Capability evals: pass@3 >= 0.90
  • Regression evals: pass^3 = 1.00 for release-critical paths

Eval Anti-Patterns

  • Overfitting prompts to known eval examples
  • Measuring only happy-path outputs
  • Ignoring cost and latency drift while chasing pass rates
  • Allowing flaky graders in release gates

Minimal Eval Artifact Layout

  • .claude/evals/<feature>.md definition
  • .claude/evals/<feature>.log run history
  • docs/releases/<version>/eval-summary.md release snapshot

版本历史

  • 581d130 当前 2026-07-19 09:07

同 Skill 集合

attic/2026-07-09/skills-pre-update/taste-skill/SKILL.md
attic/2026-07-09/skills-pre-update/ui-ux-pro-max/SKILL.md
skills/agent-development/SKILL.md
skills/api-and-interface-design/SKILL.md
skills/api-contract-standards/SKILL.md
skills/architect-system-design/SKILL.md
skills/backend-api-standards/SKILL.md
skills/backend-code-review/SKILL.md
skills/backend-error-handling/SKILL.md
skills/backend-performance-standards/SKILL.md
skills/backend-standards-always-follow/SKILL.md
skills/canary-playwright/SKILL.md
skills/caveman/SKILL.md
skills/ci-cd-and-automation/SKILL.md
skills/code-execution-standard/SKILL.md
skills/code-review-and-quality/SKILL.md
skills/code-simplification/SKILL.md
skills/codebase-design/SKILL.md
skills/codebase-start-point-guide/SKILL.md
skills/command-development/SKILL.md
skills/composition-patterns/SKILL.md
skills/context-engineering/SKILL.md
skills/dead-code-and-change-audit/SKILL.md
skills/debug-investigation/SKILL.md
skills/debugging-and-error-recovery/SKILL.md
skills/deprecation-and-migration/SKILL.md
skills/design-extract/SKILL.md
skills/design-review-playwright/SKILL.md
skills/diagnose/SKILL.md
skills/documentation-and-adrs/SKILL.md
skills/domain-modeling/SKILL.md
skills/domain-scaffold-patterns/SKILL.md
skills/doubt-driven-development/SKILL.md
skills/dox-doc-tree/SKILL.md
skills/fix-lint-format/SKILL.md
skills/forensic-change-coupling/SKILL.md
skills/forensic-complexity-trends/SKILL.md
skills/forensic-debt-quantification/SKILL.md
skills/forensic-hotspot-finder/SKILL.md
skills/frontend-api-standards/SKILL.md
skills/frontend-code-review/SKILL.md
skills/frontend-response-handling/SKILL.md
skills/frontend-server-data-patterns/SKILL.md
skills/frontend-standards-always-follow/SKILL.md
skills/frontend-structure-standards/SKILL.md
skills/frontend-ui-engineering/SKILL.md
skills/git-workflow-and-versioning/SKILL.md
skills/golang-patterns/SKILL.md
skills/golang-testing/SKILL.md
skills/graphify/SKILL.md
skills/gsd-add-tests/SKILL.md
skills/gsd-ai-integration-phase/SKILL.md
skills/gsd-audit-fix/SKILL.md
skills/gsd-audit-milestone/SKILL.md
skills/gsd-audit-uat/SKILL.md
skills/gsd-autonomous/SKILL.md
skills/gsd-capture/SKILL.md
skills/gsd-cleanup/SKILL.md
skills/gsd-code-review/SKILL.md
skills/gsd-complete-milestone/SKILL.md
skills/gsd-config/SKILL.md
skills/gsd-debug/SKILL.md
skills/gsd-discuss-phase/SKILL.md
skills/gsd-docs-update/SKILL.md
skills/gsd-eval-review/SKILL.md
skills/gsd-execute-phase/SKILL.md
skills/gsd-explore/SKILL.md
skills/gsd-extract-learnings/SKILL.md
skills/gsd-fast/SKILL.md
skills/gsd-forensics/SKILL.md
skills/gsd-graphify/SKILL.md
skills/gsd-health/SKILL.md
skills/gsd-import/SKILL.md
skills/gsd-inbox/SKILL.md
skills/gsd-ingest-docs/SKILL.md
skills/gsd-manager/SKILL.md
skills/gsd-map-codebase/SKILL.md
skills/gsd-milestone-summary/SKILL.md
skills/gsd-mvp-phase/SKILL.md
skills/gsd-new-milestone/SKILL.md
skills/gsd-new-project/SKILL.md
skills/gsd-ns-context/SKILL.md
skills/gsd-ns-ideate/SKILL.md
skills/gsd-ns-manage/SKILL.md
skills/gsd-ns-review/SKILL.md
skills/gsd-ns-workflow/SKILL.md
skills/gsd-pause-work/SKILL.md
skills/gsd-phase/SKILL.md
skills/gsd-plan-phase/SKILL.md
skills/gsd-plan-review-convergence/SKILL.md
skills/gsd-pr-branch/SKILL.md
skills/gsd-profile-user/SKILL.md
skills/gsd-progress/SKILL.md
skills/gsd-quick/SKILL.md
skills/gsd-resume-work/SKILL.md
skills/gsd-review-backlog/SKILL.md
skills/gsd-review/SKILL.md
skills/gsd-secure-phase/SKILL.md
skills/gsd-ship/SKILL.md
skills/gsd-sketch/SKILL.md
skills/gsd-spec-phase/SKILL.md
skills/gsd-spike/SKILL.md
skills/gsd-stats/SKILL.md
skills/gsd-surface/SKILL.md
skills/gsd-thread/SKILL.md
skills/gsd-ui-phase/SKILL.md
skills/gsd-ui-review/SKILL.md
skills/gsd-ultraplan-phase/SKILL.md
skills/gsd-undo/SKILL.md
skills/gsd-update/SKILL.md
skills/gsd-validate-phase/SKILL.md
skills/gsd-verify-work/SKILL.md
skills/gsd-workspace/SKILL.md
skills/gsd-workstreams/SKILL.md
skills/huashu-design/SKILL.md
skills/improve-codebase-architecture/SKILL.md
skills/incremental-implementation/SKILL.md
skills/iterative-retrieval/SKILL.md
skills/lean-ctx/SKILL.md
skills/mcp-builder/SKILL.md
skills/mcp-usage-standards/SKILL.md
skills/mmx-cli/SKILL.md
skills/owasp-security/SKILL.md
skills/pdf/SKILL.md
skills/performance-optimization/SKILL.md
skills/plan-exec-stack-guide/SKILL.md
skills/plan-mode-gate/SKILL.md
skills/planning-and-task-breakdown/SKILL.md
skills/postgres-patterns/SKILL.md
skills/project-reference-linkage/SKILL.md
skills/project-structure-map/SKILL.md
skills/qa-playwright/SKILL.md
skills/react-hooks-patterns/SKILL.md
skills/resolving-merge-conflicts/SKILL.md
skills/santa-review/SKILL.md
skills/scaffold-standards/SKILL.md
skills/security-and-hardening/SKILL.md
skills/service-layer-standards/SKILL.md
skills/shadcn/SKILL.md
skills/shipping-and-launch/SKILL.md
skills/skill-linkage-story/SKILL.md
skills/source-driven-development/SKILL.md
skills/spec-driven-development/SKILL.md
skills/strategic-compact/SKILL.md
skills/tailwind-design-system/SKILL.md
skills/taste-skill/SKILL.md
skills/tdd/SKILL.md
skills/tech-debt-audit/SKILL.md
skills/test-driven-development/SKILL.md
skills/tool-and-doc-selection/SKILL.md
skills/using-agent-skills/SKILL.md
skills/verification-loop/SKILL.md
skills/vite-react-best-practices/SKILL.md
skills/web-design-guidelines/SKILL.md
skills/webapp-testing/SKILL.md
skills/workflow-orchestrator/SKILL.md
skills/zoom-out/SKILL.md
attic/2026-07-09/skills-pre-update/huashu-design/SKILL.md
attic/2026-07-09/skills-pre-update/impeccable/SKILL.md
skills/browser-testing-with-devtools/SKILL.md
skills/codebase-intel-first/SKILL.md
skills/docx/SKILL.md
skills/find-skills/SKILL.md
skills/gsd-help/SKILL.md
skills/gsd-ns-project/SKILL.md
skills/gsd-settings/SKILL.md
skills/higgsfield-generate/SKILL.md
skills/higgsfield-marketplace-cards/SKILL.md
skills/higgsfield-product-photoshoot/SKILL.md
skills/higgsfield-soul-id/SKILL.md
skills/higgsfield-websites/SKILL.md
skills/impeccable/SKILL.md
skills/jcodemunch-token-saver/SKILL.md
skills/pptx/SKILL.md
skills/tdd-auto-init/SKILL.md
skills/ui-ux-pro-max/SKILL.md
skills/update-docs/SKILL.md
skills/xlsx/SKILL.md
skills/autoplan/SKILL.md
skills/benchmark/SKILL.md
skills/browse/SKILL.md
skills/careful/SKILL.md
skills/connect-chrome/SKILL.md
skills/context-restore/SKILL.md
skills/design-consultation/SKILL.md
skills/design-html/SKILL.md
skills/design-review/SKILL.md
skills/design-shotgun/SKILL.md
skills/diagram/SKILL.md
skills/document-generate/SKILL.md
skills/freeze/SKILL.md
skills/guard/SKILL.md
skills/investigate/SKILL.md
skills/ios-clean/SKILL.md
skills/ios-design-review/SKILL.md
skills/ios-sync/SKILL.md
skills/landing-report/SKILL.md
skills/make-pdf/SKILL.md
skills/open-gstack-browser/SKILL.md
skills/pair-agent/SKILL.md
skills/plan-design-review/SKILL.md
skills/plan-devex-review/SKILL.md
skills/plan-tune/SKILL.md
skills/qa/SKILL.md
skills/setup-browser-cookies/SKILL.md
skills/setup-deploy/SKILL.md
skills/setup-gbrain/SKILL.md
skills/ship/SKILL.md
skills/skillify/SKILL.md
skills/spec/SKILL.md
skills/sync-gbrain/SKILL.md
skills/unfreeze/SKILL.md
skills/benchmark-models/SKILL.md
skills/canary/SKILL.md
skills/codex/SKILL.md
skills/context-save/SKILL.md
skills/cso/SKILL.md
skills/devex-review/SKILL.md
skills/document-release/SKILL.md
skills/gstack-upgrade/SKILL.md
skills/health/SKILL.md
skills/ios-fix/SKILL.md
skills/ios-qa/SKILL.md
skills/land-and-deploy/SKILL.md
skills/learn/SKILL.md
skills/office-hours/SKILL.md
skills/plan-ceo-review/SKILL.md
skills/plan-eng-review/SKILL.md
skills/qa-only/SKILL.md
skills/retro/SKILL.md
skills/review/SKILL.md
skills/scrape/SKILL.md

元信息

文件数
0
版本
581d130
Hash
b29e7054
收录时间
2026-07-19 09:07

首页 - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-07-22 02:48
浙ICP备14020137号-1 $访客地图$