Agent Skills
› Gentleman-Programming/gentle-ai
› gentle-ai-bench
gentle-ai-bench
GitHub用于 gentle-ai 项目的基准测试旅程编写与验证。指导如何创建、审查 journey 用例,确保 ID 唯一性及 Review 声明合规。强调通过构建二进制并运行 harness 来提供驱动执行的真实证据,而非仅依赖声明检查。
Trigger Scenarios
修改 bench/ 目录下的 journey 文件
产品语义变更需更新 pinned journey
诊断 CI Unit Tests job 中的 bench 失败
Install
npx skills add Gentleman-Programming/gentle-ai --skill gentle-ai-bench -g -y
SKILL.md
Frontmatter
{
"name": "gentle-ai-bench",
"license": "Apache-2.0",
"metadata": {
"author": "Gentleman-Programming",
"version": "1.0"
},
"description": "Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis. Author and verify gentle-ai bench journeys; go test .\/bench never proves driven execution."
}
Activation Contract
Load when touching bench/ in gentle-ai, adding or changing a journey, changing a product semantic a journey might pin, or diagnosing a bench failure in CI's Unit Tests job.
Hard Rules
go test ./benchvalidates corpus declarations only. It does NOT execute journeys. The only driven proof is building the harness and the product binary and running the harness against it; a greengo test ./benchclaims nothing about execution.- Reproduce CI, do not guess invocations: read the Unit Tests step in
.github/workflows/ci.ymland copy its exact build andgentle-ai-bench run --binary ...commands. Use--only <journey-id>to drive one journey. - Journey IDs are unique across every
journeys_*.gofile. The collision guard fails loudly naming both files; pick an unused ID by reading the corpus, never reuse a retired one. - Every journey declares
Review:—reviewOptedIn(the runner enables receipt-driven development globally before the first step, uncounted, and fails the journey if the switch does not come on) orreviewUntouched(its subject IS the switch, or it has nothing to do with reviews). The declaration is mandatory;validateCorpusfails the run without it. Never let a journey inherit the product's default: reviews are opt-in, and a journey that assumed otherwise measures a review-refused flow while still reportingcompleted. - Every
executetransition must carry a runnable command; the dead-execute guard fails the run otherwise. - When a ratified product semantic changes, grep the corpus for journeys pinning the OLD behavior before shipping. The corpus is a second test surface beyond unit tests; a journey asserting the defect keeps the defect green.
dead_endprintsn/aunless the run actually measured one. Never fabricate a value to move the column.- A
by_designexemption costs a shape from the closed vocabulary plus a verified quote of the product's own next-action text. If the quote no longer tells the operator what to do, it is a defect wearing an exemption. - Prefer a NEW
journeys_*.gofile when the shared ones are owned by open PRs; bump the core journey-count pin in the same change.
Execution Steps
- Read the corpus area you touch and the CI invocation before writing.
- Author or adapt the journey; update its title, step names, and comment to say WHY the expectation holds (cite the issue or ratified decision).
- Run
go test ./...inbench/for declarations, THEN the driven harness for execution; both results go in the PR body. - On semantic changes, list the journeys you checked for stale pins.
Output Contract
PR evidence includes the driven-mode summary line (completed / unsupported / failed counts) from a locally built binary, not only go test output.
Version History
- 35deba3 Current 2026-08-20 00:48


