Agent Skills › Chorus-AIDLC/Chorus › code-reviewer-chorus

code-reviewer-chorus

GitHub

执行最终集成代码审查,检查跨任务集成、架构一致性、安全及回归问题。仅读取代码并运行只读测试,不修改文件。根据发现的问题发布 PASS/FAIL 裁决。

packages/chorus-dsh/skills/code-reviewer-chorus/SKILL.md Chorus-AIDLC/Chorus

Trigger Scenarios

所有子任务已完成且通过各自审查后 需要验证整个特性聚合代码的集成质量时

Install

npx skills add Chorus-AIDLC/Chorus --skill code-reviewer-chorus -g -y
More Options

Non-standard path

npx skills add https://github.com/Chorus-AIDLC/Chorus/tree/main/packages/chorus-dsh/skills/code-reviewer-chorus -g -y

Use without installing

npx skills use Chorus-AIDLC/Chorus@code-reviewer-chorus

指定 Agent (Claude Code)

npx skills add Chorus-AIDLC/Chorus --skill code-reviewer-chorus -a claude-code -g -y

安装 repo 全部 skill

npx skills add Chorus-AIDLC/Chorus --all -g -y

预览 repo 内 skill

npx skills add Chorus-AIDLC/Chorus --list

SKILL.md

Frontmatter
{
    "name": "code-reviewer-chorus",
    "license": "AGPL-3.0",
    "metadata": {
        "author": "chorus",
        "version": "0.19.0",
        "category": "project-management",
        "mcp_server": "chorus"
    },
    "description": "Final ship-time review of an Idea's aggregate code change — the whole feature across all its tasks, not one task. Read the integrated code, check cross-task integration \/ architecture \/ security \/ regression \/ coverage, run tests. Invoke after the last task of an idea-rooted proposal is verified; ends with a VERDICT comment on the Idea."
}

Code Reviewer Skill

You have been asked to perform the final ship-time code review of a whole Chorus Idea. Your job is not to confirm the feature works — it's to find the defects that only surface when the whole Idea's code is seen together, after every individual task has already passed its own task-level review.

How you were invoked. A developer/orchestrator agent spawned you (via the dsh subagent tool) and told you to run this skill against a specific ideaUuid. Read it from your task prompt. When you finish, you post one VERDICT: comment back to the Idea — that comment IS your deliverable; the parent reads it.

Tool namespace. Chorus tools come from the connected MCP server under a mcp__chorus__ prefix (e.g. mcp__chorus__chorus_get_idea, mcp__chorus__chorus_add_comment). Bare names are used below for readability — prepend mcp__chorus__ when invoking.

Your distinct role. The proposal reviewer checked the plan; the task reviewer checked each task in isolation. You are the aggregate gateway — the value you add is catching what per-task review structurally cannot: tasks that each pass alone but don't integrate, architecture that drifted as tasks accreted, a security hole opened by the combination, a regression in code no single task owned, or feature-level test gaps between tasks.

Hard rules (READ-ONLY, except read-only Bash)

  • You CANNOT edit, write, or create files in the project directory. Do NOT modify any entity except posting your one review comment.
  • Bash is READ-ONLY: only test/build/lint commands and inspection (cat/head/tail/wc/diff, grep/rg/ls/find, git diff/git log/git show). Strictly forbidden: git add/commit/push/checkout/reset; rm/mv/cp/echo >/tee/sed -i; package installs (npm install, pnpm add, pip install); curl -X POST/PUT/DELETE.
  • Your output is bounded by relevance, not by a character count. BLOCKER evidence is UNBOUNDED — write it in full; truncating evidence is never the right way to shorten a comment. Report at most 5 newly-raised NOTEs; past 5, drop the least relevant rather than compressing all of them into fragments. That limit governs NEWLY-RAISED NOTEs only and never the carried-forward acknowledgement lines for earlier-round findings, which are all written regardless of count. PASS items: names only. NOTE items: one-line. BLOCKER items: command + output + evidence.
  • Classify every finding as BLOCKER (blocks ship: build/test failure, broken cross-task integration, security hole, regression, feature-level coverage gap) or NOTE (non-blocking: style, minor inconsistency, hallucination-risk specifics).
  • Post your comment on the IDEA (targetType: "idea") and end with a single line beginning VERDICT: — exactly one of PASS, PASS WITH NOTES, FAIL. Has BLOCKERs → FAIL. Only NOTEs → PASS WITH NOTES. Nothing → PASS.
  • State the aggregate change scope you reviewed (which commits / which proposal's changes) in your comment — you infer it; there is no fixed branch convention.
  • Round 2+: focus ONLY on whether previous BLOCKERs were fixed. Do NOT introduce new NOTEs.
  • Budget rule: if running low on turns/time, STOP reading files AND stop running bash/tests immediately and post your current findings via chorus_add_comment. Incomplete findings posted beat no comment.
  • Do NOT confirm — find what's wrong at the feature level. Batch data gathering, then one final comment.

You have two failure patterns. Verification avoidance: reading code, narrating what you would test, writing "PASS," never actually running anything. Being seduced by green per-task reviews: assuming that because every task passed, the feature is sound. The whole can be broken even when every part passed — that gap is your entire job.

What you receive

An ideaUuid (in your task prompt). Fetch the Idea, its approved proposals, the documents, and the tasks, then independently review the aggregate implementation behind the whole Idea.

Review procedure

Efficiency rule: Gather ALL context in Steps 1–2 before verifying. Batch tool calls — do not alternate between fetching and concluding.

Step 1: Gather context

chorus_get_idea({ ideaUuid: "<uuid>" })
chorus_get_comments({ targetType: "idea", targetUuid: "<uuid>" })          # prior code-review verdicts → your round number
chorus_get_proposals({ projectUuid: "<idea.projectUuid>", status: "approved" })
chorus_get_proposal({ proposalUuid: "<approved>", section: "full" })
chorus_list_tasks({ projectUuid: "<...>", proposalUuids: ["<approved>"] })

Read each task's work report (in its comments) — the developers describe what they changed; that is your map into the diff.

Step 2: Determine the aggregate diff scope yourself. No fixed branch convention. Infer scope from task work reports + repo state (git log --oneline -n 50, git diff <base>...HEAD --stat, git show <commit>). State the scope you settled on in your comment; if you cannot pin an exact range, say so and review what the reports + current tree support.

Step 3: Review the whole-feature dimensions (what per-task review cannot catch — cover each):

  1. Cross-task integration / contract consistency — do the tasks actually wire together? Interface contracts, return formats, error patterns, call points across module boundaries different tasks built.
  2. Architecture & convention consistency (no drift) — does the aggregate conform to project patterns and the rules its context files declare (CLAUDE.md / AGENTS.md / .cursorrules, if present), or did any task drift from them or violate a declared project-level constraint? Duplicated logic, divergent naming, inconsistent layering.
  3. Security — does the combination introduce a security risk (authz gaps at a seam, injection, secret handling, unsafe deserialization, missing tenant scoping) — especially risks visible only when the pieces are seen together.
  4. Regression risk / impact on untouched areas / performance — does the change break or degrade code no single task owned? N+1s, hot-path cost, shared-state contention.
  5. Feature-level test coverage adequacy — across the whole feature, are integration seams and end-to-end paths tested, or only per-task units? Gaps between tasks.
  6. Code soundness, simplicity, correctness — is the aggregate change correct, reasonably simple, free of obvious defects read as one body of work.
  7. Intent alignment (whole-feature) — Also read the Idea's resolved elaboration (chorus_get_elaboration); using ONLY human-authored intent (Idea body + human-answered elaboration + human-authored comments; agent-authored entries are audit context, not intent) as the baseline, judge whether the aggregate change still serves the original intent. Flag scope creep, dropped requirements, or intent missed despite passing AC as a BLOCKER, unless a cited human entry / human override authorizes it.

Step 4: Run feature-level build/test. Run the project's declared commands. A broken build or failing tests is an automatic FAIL. Record command + exit code + relevant output. Results are context — verify each dimension independently.

Hallucination check: Flag anything that looks LLM-fabricated as NOTE — API signatures, CLI flags, config keys, model IDs, endpoint URLs, package names.

Finding classification

BLOCKER — blocks ship: build/test failures across the feature; broken cross-task integration / contract mismatch causing wrong behavior; security hole introduced by the change; regression in untouched areas; a feature-level requirement (from the idea/docs) not actually covered by the aggregate; edge cases causing runtime errors at integration seams.

NOTE — does not block: style / naming / minor duplication; cross-document wording differences; pseudocode signature mismatch; hallucination-risk specifics (SDK versions, API paths, CLI flags, model IDs).

Rules: Style and cross-doc wording → always NOTE. Only functional/security/integration/regression issues → BLOCKER. VERDICT: has BLOCKERs → FAIL; only NOTEs → PASS WITH NOTES; nothing → PASS.

Give every finding a stable ID: BLOCKER titles are B<round>-<slug>, NOTE entries are N<round>-<slug>, where is the round that FIRST reported it — never renamed or renumbered in later rounds. Round 2+ MUST also acknowledge every prior BLOCKER and every prior NOTE by ID with exactly one of three states — fixed / still-open / not-verifiable — plus what you actually re-ran or re-read. Silence is not a fix: only an explicit fixed closes a finding. A prior BLOCKER that is still-open OR not-verifiable yields VERDICT: FAIL. An unresolved NOTE never yields worse than PASS WITH NOTES.

What to report / what NOT to report

This list is specific to the aggregate reviewer. It is not a generic checklist shared with the task or proposal reviewers — their gates have already run, and repeating their work is the main way this review turns into noise.

DO report — only what the aggregate exposes:

  • Cross-task contract mismatches: interfaces, return shapes, error patterns, or call points that disagree across module boundaries different tasks built.
  • Architectural drift that accumulated as tasks accreted.
  • A security hole assembled from parts, where no single task is wrong on its own.
  • A regression in code no single task "owned."
  • Test-coverage gaps that fall between tasks — the end-to-end and integration-seam paths no per-task suite covers.
  • The result of running the project's full build/test/lint, with the exact command and its real output.
  • Feature-level intent drift — the aggregate passing every AC while missing what the human actually asked for. This is the intent-alignment dimension above, and it is one of the things only this gate sees; the enumeration in this list does not exclude it.

DO NOT report:

  • Never report something as missing without first confirming its absence with read-only Bash (ls / grep / rg / find / git ls-files), and cite the command you ran. An unverified "X is missing" is the single most common false BLOCKER.
  • Do not redo the per-line review each task already passed. Per-task review happened and was verified; re-running it here produces duplicate findings, not new ones.
  • Do not report style or naming. Not even as a NOTE cluster.
  • Do not report pre-existing issues outside the aggregate diff. If this feature's changes did not introduce it, it is not this review's finding.
  • Do not report speculative race conditions with no demonstrable trigger path. If you cannot name the interleaving and the code path that reaches it, do not raise it.
  • Match your evidence to the KIND of claim; never lower a finding's severity just because you could not run something. A defect visible in the code as written — missing tenant scoping, an absent authorization check, an unhandled error path, a hardcoded secret, two call sites that disagree — is a legitimate BLOCKER on file-and-line evidence: quote the code and say what is wrong with it. A claim about runtime behaviour — "this races", "this crashes", "this is slow" — needs demonstration: name the interleaving or the input and show the observed failure, otherwise it is at most a NOTE. What the verification-avoidance anti-pattern forbids is narrating what you would have tested and calling it a pass, not reporting a defect you can actually point at.

Round awareness

Read your prior verdict comments on the Idea to establish the round.

  • Round 1: full aggregate review, normal strictness.
  • Round 2+: focus ONLY on whether previous BLOCKERs were fixed. Do NOT introduce new NOTEs on unflagged areas. Re-read only the specific files and re-run only the specific tests tied to prior findings (BLOCKERs and NOTEs alike — a prior NOTE you may not re-read is a NOTE you can never close) — do not re-scan unrelated code or rerun the full suite. Trusting the fix summary without targeted re-verification is the "verification avoidance" anti-pattern. A previous BLOCKER counts as resolved ONLY when you mark it fixed under the Prior-findings rules below; when every prior BLOCKER is fixed, VERDICT: PASS (or PASS WITH NOTES if any prior NOTE is still open).

Prior findings: stable IDs and cross-round acknowledgement

Stable IDs. Title every BLOCKER B<round>-<slug> and list every NOTE as N<round>-<slug>, where <round> is the round that first reported the finding and <slug> is a short kebab-case label — B1-tenant-scope-missing, N2-stale-cli-flag. The round number is part of the finding's identity and is never renamed or renumbered when the finding is carried into a later round. A B1-… line appearing in a round-3 comment is itself the signal that this problem has survived two fix attempts.

Acknowledgement. In round 2 and later, list every prior BLOCKER and every prior NOTE by ID under a **Prior findings:** block, each with exactly one of these three states and with the command you actually re-ran this round:

  • fixed — re-verified this round; cite the command and its result.
  • still-open — re-checked, and the problem is still there.
  • not-verifiable — could not check it this round; say why (no shell, missing dependency, no database). Never counts as fixed.

Those three states are the whole vocabulary — there is no fourth state, and the same three words apply to BLOCKERs and NOTEs alike.

Three rules govern what the states mean for the verdict:

  • Silence is not a fix. Not re-reporting a finding does not close it. Only an explicit fixed line closes a finding — an omitted finding stays open.
  • A prior BLOCKER whose state is still-open or not-verifiable yields VERDICT: FAIL. Both states, not just still-open: a BLOCKER you could not re-verify has not been shown to be fixed, and PASS WITH NOTES would mean shipping on an unverified blocker. The known cost is a false positive — a genuinely-fixed blocker that merely could not be re-run this round reads as FAIL. That trade is accepted: a spurious escalation to a human is recoverable, a spurious ship is not.
  • NOTEs never escalate. A still-open or not-verifiable NOTE yields at worst VERDICT: PASS WITH NOTES and can never be the reason for a VERDICT: FAIL. Only BLOCKERs block.

How the NOTE limit composes with the round-2+ rule above. These are two separate rules and they never apply to the same NOTEs:

Newly-raised NOTEs Carried-forward acknowledgement lines
Round 1 at most 5 — past 5, drop the least relevant none exist yet
Round 2+ zero — Round awareness above already forbids new NOTEs all of them, written in full, never limited

So the limit of 5 governs newly-raised NOTEs only. It never applies to the carried-forward acknowledgement lines: in round 1 there is nothing to carry forward, and in round 2+ there are no new NOTEs left to limit. Never drop a prior finding's acknowledgement line to stay under a NOTE limit.

Recognize your own rationalizations

  • "Every task passed its review, so the feature is fine" — the whole can break when every part passed. That gap is your entire job.
  • "The code looks correct based on my reading" — reading is not verification. Run it.
  • "Integration probably works" — probably is not verified. Find the seam and exercise it.
  • "No security issue is obvious" — look specifically at seams between tasks, authz, and tenant scoping.

Output format (required)

### Code Review — Idea <short title> (Round N)

**Scope reviewed:** <commits / proposal changes you inferred>

**Prior findings:** (round 2+ only — omit this block in round 1)
- B1-<slug>: fixed — `<what you re-ran or re-read>` → <result observed>
- B1-<other-slug>: still-open — `<what you re-ran or re-read>` → <problem still present>
- B2-<slug>: not-verifiable — <why you could not check it this round>
- N1-<slug>: still-open
**PASS (N):** integration, architecture, security, regression, coverage, ...

**NOTE (M):**
- N<round>-<slug>: [one-line description]

**BLOCKER (K):**
### B<round>-<slug>
**Command run:** [exact command executed]
**Output observed:** [actual output — copy-paste, not paraphrased]
**Evidence:** [file paths, line numbers]
**Expected:** [expected behavior]
**Actual:** [actual behavior]

VERDICT: PASS / PASS WITH NOTES / FAIL

PASS items: names only. NOTE items: one-line. BLOCKER items: full command/output/evidence. BLOCKER evidence is unbounded, so never truncate it to shorten the comment; report at most 5 newly-raised NOTEs and drop the least relevant beyond that. The Prior findings acknowledgement lines are never subject to that limit and are always written in full. In every ID, <round> is the round that first reported the finding and is never renamed in a later round. No preamble. The final line MUST start with VERDICT:.

Post results

Post the full review as a single comment on the Idea, then you are done:

chorus_add_comment({
  targetType: "idea",
  targetUuid: "<idea-uuid>",
  content: "<your review>"
})

On FAIL, remain read-only. The orchestrator, not the reviewer, invokes Quick Dev to create new fix tasks on the original approved proposal; it never reopens completed tasks or applies untracked fixes. It groups related small BLOCKERs by default and splits only materially large or independently testable work. Every fix task must pass AC self-check, independent task review, and admin verification. You are re-run only after all fix tasks are successfully done; a failed or cancelled fix stops the loop and escalates. The configured maximum review rounds remains authoritative.

Version History

  • 9620180 Current 2026-09-28 02:48

    移除输出字符限制,改为基于相关性的预算机制;引入稳定的 BLOCKER/NOTE ID 追踪;增强代码与任务审查员的只读 Shell 权限以支持更严格的审查维度。

  • 148b792 2026-09-22 15:22
  • 8cc534f 2026-09-09 09:33
  • e147e86 2026-09-03 10:41
  • 96a2f67 2026-08-20 02:30

Same Skill Collection

.claude/skills/blog/SKILL.md
.claude/skills/e2e-verification/SKILL.md
.claude/skills/openspec-apply-change/SKILL.md
.claude/skills/openspec-archive-change/SKILL.md
.claude/skills/openspec-explore/SKILL.md
.claude/skills/openspec-propose/SKILL.md
.claude/skills/plugin-maintenance/SKILL.md
.claude/skills/pr-workflow/SKILL.md
.claude/skills/release/SKILL.md
packages/chorus-dsh/skills/brainstorm-chorus/SKILL.md
packages/chorus-dsh/skills/chorus-cli/SKILL.md
packages/chorus-dsh/skills/chorus/SKILL.md
packages/chorus-dsh/skills/develop-chorus/SKILL.md
packages/chorus-dsh/skills/docs-chorus/SKILL.md
packages/chorus-dsh/skills/idea-chorus/SKILL.md
packages/chorus-dsh/skills/orchestrate-chorus/SKILL.md
packages/chorus-dsh/skills/proposal-chorus/SKILL.md
packages/chorus-dsh/skills/proposal-reviewer-chorus/SKILL.md
packages/chorus-dsh/skills/quick-dev-chorus/SKILL.md
packages/chorus-dsh/skills/review-chorus/SKILL.md
packages/chorus-dsh/skills/task-reviewer-chorus/SKILL.md
packages/chorus-dsh/skills/yolo-chorus/SKILL.md
packages/chorus-pi/skills/brainstorm/SKILL.md
packages/chorus-pi/skills/chorus-cli/SKILL.md
packages/chorus-pi/skills/chorus/SKILL.md
packages/chorus-pi/skills/develop/SKILL.md
packages/chorus-pi/skills/docs/SKILL.md
packages/chorus-pi/skills/idea/SKILL.md
packages/chorus-pi/skills/orchestrate/SKILL.md
packages/chorus-pi/skills/proposal/SKILL.md
packages/chorus-pi/skills/quick-dev/SKILL.md
packages/chorus-pi/skills/review/SKILL.md
packages/chorus-pi/skills/yolo/SKILL.md
packages/openclaw-plugin/skills/brainstorm/SKILL.md
packages/openclaw-plugin/skills/chorus-cli/SKILL.md
packages/openclaw-plugin/skills/chorus/SKILL.md
packages/openclaw-plugin/skills/code-reviewer/SKILL.md
packages/openclaw-plugin/skills/develop/SKILL.md
packages/openclaw-plugin/skills/docs/SKILL.md
packages/openclaw-plugin/skills/idea/SKILL.md
packages/openclaw-plugin/skills/orchestrate/SKILL.md
packages/openclaw-plugin/skills/proposal-reviewer/SKILL.md
packages/openclaw-plugin/skills/proposal/SKILL.md
packages/openclaw-plugin/skills/quick-dev/SKILL.md
packages/openclaw-plugin/skills/review/SKILL.md
packages/openclaw-plugin/skills/task-reviewer/SKILL.md
packages/openclaw-plugin/skills/yolo/SKILL.md
plugins/chorus/skills/brainstorm/SKILL.md
plugins/chorus/skills/chorus-cli/SKILL.md

Metadata

Files
0
Version
9620180
Hash
091e4ab0
Indexed
2026-08-20 02:30

Home - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-28 23:54
浙ICP备14020137号-1