Agent Skillsmohitagw15856/pm-claude-skills › agent-design-review

agent-design-review

GitHub

审查LLM智能体架构,识别不可靠、高成本或安全隐患。覆盖控制流、工具使用、记忆、失败处理及安全性等维度,输出结构化报告、风险评估与优先级修复方案,提升生产环境稳定性。

exports/openclaw/agent-design-review/SKILL.md mohitagw15856/pm-claude-skills

触发场景

审查智能体架构设计 调试陷入循环或偏离任务的智能体 发布前加固智能体安全性 评估多步/工具调用智能体的可靠性

安装

npx skills add mohitagw15856/pm-claude-skills --skill agent-design-review -g -y
更多选项

非标准路径

npx skills add https://github.com/mohitagw15856/pm-claude-skills/tree/main/exports/openclaw/agent-design-review -g -y

不安装直接使用

npx skills use mohitagw15856/pm-claude-skills@agent-design-review

指定 Agent (Claude Code)

npx skills add mohitagw15856/pm-claude-skills --skill agent-design-review -a claude-code -g -y

安装 repo 全部 skill

npx skills add mohitagw15856/pm-claude-skills --all -g -y

预览 repo 内 skill

npx skills add mohitagw15856/pm-claude-skills --list

SKILL.md

Frontmatter
{
    "name": "agent-design-review",
    "homepage": "https:\/\/mohitagw15856.github.io\/pm-claude-skills\/skill\/agent-design-review.html",
    "metadata": {
        "openclaw": {
            "emoji": "🤖"
        }
    },
    "description": "Review an LLM agent design and find where it will be unreliable, expensive, or unsafe. Use when asked to review an agent architecture, critique a multi-step\/tool-using agent, debug an agent that loops or goes off-task, or harden an agent before launch. Produces a structured review — task fit, control flow, tools, memory\/context, failure handling, cost, and safety — with prioritised findings and fixes."
}

Agent Design Review Skill

Most agents don't fail because the model is weak — they fail because the design lets them loop, call the wrong tool, lose the thread across steps, or burn tokens with no stopping rule. This skill reviews an agent's architecture against the decisions that actually determine reliability, and ranks the fixes — so "it works in the demo but not in prod" becomes a specific list of changes. (Writing a new agent spec? Use agent-spec.)

Working from a brief

Given a sketch ("a research agent that searches, reads, and writes a report"), deliver the full review anyway — infer the likely control flow and tools, label the inference, and flag what to confirm. Never withhold the review for missing detail.

Required Inputs

Ask for these only if they aren't already provided (else infer and label):

  • What the agent does — its goal, and what a successful run produces.
  • Control flow — single prompt, plan-then-execute, ReAct loop, or multi-agent; and the stopping condition.
  • Tools & actions — what it can call, and which actions have side effects (write, send, pay).
  • Memory & context — what state carries across steps, and how context is kept in budget.
  • Constraints — latency, cost per run, and the trust boundary (untrusted input? real-world actions?).

Output Format

Agent Review: [agent]

1. Summary — will this be reliable in production? The top 3 risks and the single change that helps most.

2. Findings by dimension — for each, what's sound and what's fragile:

Dimension Finding Severity Fix
Control flow no max-steps / no progress check → loops High step budget + "am I making progress?" check + halt
Tool use overlapping tools confuse selection Med fewer, sharply-described tools; allowlist
Context full history re-sent each step → cost + drift High summarise/scope memory per step
Failure handling one tool error aborts the run Med retry/backoff + graceful degradation
Safety acts without confirmation on writes High human/confirm gate on side-effecting actions

3. Reliability checklist — termination guarantee (it always stops), error recovery, idempotency of side-effecting actions, and determinism where it matters.

4. Cost & latency — where tokens/steps are spent and how to cut them (cheaper model for sub-steps, caching, fewer round-trips) without losing quality. Pair with llm-cost-latency-budget.

5. Safety — untrusted input/tool output handled as data not instructions, least-privilege tools, and confirmation gates on high-impact actions. Pair with llm-guardrails-spec.

6. Prioritised fix plan — ordered by impact-to-effort.

Quality Checks

  • The agent has a guaranteed stopping condition (step/budget cap + progress check) — no unbounded loops
  • Side-effecting actions are idempotent or gated by a confirmation
  • Tools are few and sharply described so selection is unambiguous; access is least-privilege
  • Context strategy keeps the window in budget across steps (no naive full-history resend)
  • Tool errors are recovered, not fatal — retry/backoff and graceful degradation
  • Findings are severity-ranked and the fix plan is ordered by impact

Anti-Patterns

  • Do not approve an agent with no termination guarantee — "it usually stops" is an outage waiting to happen
  • Do not let it take irreversible actions without a confirmation gate
  • Do not give it many overlapping tools — selection accuracy drops as the toolset grows
  • Do not resend the whole history every step — cost and drift both climb
  • Do not treat tool/retrieved output as trusted instructions — it's the injection surface

Based On

LLM agent design practice — bounded control flow, least-privilege tool use, context management, error recovery, and safety gating.

版本历史

  • 54fad50 当前 2026-07-19 12:09

同 Skill 集合

exports/openclaw/360-feedback-template/SKILL.md
exports/openclaw/401k-plan-decoder/SKILL.md
exports/openclaw/ab-test-planner/SKILL.md
exports/openclaw/ab-test-readout/SKILL.md
exports/openclaw/accessibility-audit/SKILL.md
exports/openclaw/account-plan/SKILL.md
exports/openclaw/acquirer-red-team/SKILL.md
exports/openclaw/ad-copy/SKILL.md
exports/openclaw/aeo-optimizer/SKILL.md
exports/openclaw/agenda-or-cancel/SKILL.md
exports/openclaw/agent-hiring-panel/SKILL.md
exports/openclaw/agent-observability-spec/SKILL.md
exports/openclaw/agent-severance/SKILL.md
exports/openclaw/agent-spec/SKILL.md
exports/openclaw/agm-in-a-box/SKILL.md
exports/openclaw/ai-ethics-review/SKILL.md
exports/openclaw/ai-eval-plan/SKILL.md
exports/openclaw/ai-feature-prd/SKILL.md
exports/openclaw/ai-product-canvas/SKILL.md
exports/openclaw/air-quality/SKILL.md
exports/openclaw/altitude-shifter/SKILL.md
exports/openclaw/ambiguity-resolver/SKILL.md
exports/openclaw/analyst-relations-brief/SKILL.md
exports/openclaw/announcement-card/SKILL.md
exports/openclaw/api-docs-writer/SKILL.md
exports/openclaw/api-test-plan/SKILL.md
exports/openclaw/api-versioning-strategy/SKILL.md
exports/openclaw/apology-letter/SKILL.md
exports/openclaw/architecture-decision-record/SKILL.md
exports/openclaw/architecture-diagram/SKILL.md
exports/openclaw/archive-strategy/SKILL.md
exports/openclaw/assumption-bounty/SKILL.md
exports/openclaw/assumption-mapper/SKILL.md
exports/openclaw/async-update-format/SKILL.md
exports/openclaw/auto-repair-estimate-decoder/SKILL.md
exports/openclaw/autopilot-charter/SKILL.md
exports/openclaw/behavior-intervention-plan/SKILL.md
exports/openclaw/benefits-decoder/SKILL.md
exports/openclaw/bennett-time-audit/SKILL.md
exports/openclaw/bid-tender-review/SKILL.md
exports/openclaw/board-deck-narrative/SKILL.md
exports/openclaw/board-game-designer/SKILL.md
exports/openclaw/board-minutes/SKILL.md
exports/openclaw/board-pre-read/SKILL.md
exports/openclaw/bom-cost-review/SKILL.md
exports/openclaw/bookkeeping-categorization/SKILL.md
exports/openclaw/boolean-search-builder/SKILL.md
exports/openclaw/brag-doc/SKILL.md
exports/openclaw/brainstorming/SKILL.md

元信息

文件数
0
版本
e4def4c
Hash
23f51544
收录时间
2026-07-19 12:09

首页 - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-08-01 08:39
浙ICP备14020137号-1 $访客地图$