Agent Skillsmohitagw15856/pm-claude-skills › ai-assisted-performance-review

ai-assisted-performance-review

GitHub

评估AI辅助下的绩效,区分人与工具贡献。提供指标重写、混合团队校准规则及对话脚本,聚焦判断力、验证等核心人类价值,避免以量代质。

skills/ai-assisted-performance-review/SKILL.md mohitagw15856/pm-claude-skills

Trigger Scenarios

员工工作大量由AI辅助 产出量不再代表努力程度 团队AI采用率不均需校准 制定AI时代评审标准

Install

npx skills add mohitagw15856/pm-claude-skills --skill ai-assisted-performance-review -g -y
More Options

Use without installing

npx skills use mohitagw15856/pm-claude-skills@ai-assisted-performance-review

指定 Agent (Claude Code)

npx skills add mohitagw15856/pm-claude-skills --skill ai-assisted-performance-review -a claude-code -g -y

安装 repo 全部 skill

npx skills add mohitagw15856/pm-claude-skills --all -g -y

预览 repo 内 skill

npx skills add mohitagw15856/pm-claude-skills --list

SKILL.md

Frontmatter
{
    "name": "ai-assisted-performance-review",
    "description": "Evaluate performance fairly when output is AI-assisted — what still measures the human, what now measures the tooling, and how to run the review conversation. Use when reviewing someone whose work is heavily AI-assisted, when output volume stopped meaning anything, when calibrating a team with uneven AI adoption, or when writing review criteria for the AI era. Produces review guidance: a what-measures-whom analysis, rewritten criteria, calibration rules for mixed-adoption teams, and conversation scripts. For the general review document use performance-review; for redesigning the role itself use role-redesign-for-ai."
}

AI-Assisted Performance Review Skill

The uncomfortable review question of the decade: when a report ships twice the output with AI, what did they do? Volume stopped measuring effort; polish stopped measuring skill. Punishing AI use is as wrong as crediting the model's work to the human. This skill separates the signals — and gives managers the conversation, not just the theory.

What This Skill Produces

  • A what-measures-whom analysis of the role's current evaluation criteria
  • Rewritten criteria that measure the human: judgment, verification, outcomes, leverage
  • Calibration rules for teams with uneven AI adoption
  • Conversation scripts for the three hard cases

Required Inputs

Ask for (if not already provided):

  • The role and current review criteria (the rubric, or how it really works)
  • How AI shows up in the work — which tasks, how much of the output it drafts, what the tooling reality is
  • The specific situation, if any: one person's review? team calibration? criteria rewrite?
  • The org's AI stance — encouraged? tolerated? policy exists? (Reviews must not punish sanctioned behaviour)

Method

  1. Sort every criterion: human, tool, or hybrid. Walk the current rubric. Volume of drafts, formatting quality, speed to first version → now mostly tool signals (evaluating them evaluates prompt luck and subscription tier). Decision quality, stakeholder trust, error catch rate, what they chose to build → still human. Output quality overall → hybrid: credit belongs to the pair, and the review's job is to see the human's contribution inside it.
  2. Rewrite around the four durable human signals:
    • Judgment — what they decided to do, what they declined, how they scoped; the quality of taste applied to AI output (what they kept, cut, and corrected)
    • Verification — do errors get caught before shipping? A person whose AI-assisted work is reliably right is demonstrating skill; one who forwards unverified fluency is a risk wearing productivity's clothes
    • Outcomes — did the work move what it was for (the metric, the decision, the customer), independent of how it was produced
    • Leverage — do they make AI multiply the team (shared prompts, workflows, teaching) or only their own count
  3. Set the calibration rules for mixed adoption. In one team you'll have a 2×-output adopter and a careful non-adopter. Rules that keep it fair: evaluate against the role's outcomes, not each other's volume · where AI use is sanctioned, not adopting is a development conversation (not a values one) · where someone's edge is invisible verification labour, surface it explicitly before comparing. Never let the review become a proxy war about the tools.
  4. Demand evidence that sees the human. Volume anecdotes are out. In: a sample of shipped work walked backwards (what did the AI draft, what did you change, why) · error/rework history · decisions log · peer signals about trust and leverage. The walk-backwards exercise is the single highest-signal artifact — put it in the review prep.
  5. Script the three hard cases:
    • The volume star with thin judgment — "Your output doubled; let's walk three pieces backwards" (the conversation is about the delta between draft and shipped)
    • The careful sceptic being out-shipped — outcomes-first framing; adoption raised as growth, not deficiency; their verification strength named as a strength
    • The launderer — unverified AI work shipped as their own, errors reaching others: this is a reliability conversation with the accountability rule from the org's AI policy, not an AI conversation

Output Format

AI-Era Review Guidance: [role/team]

Criteria audit

Current criterion Measures Verdict
human / tool / hybrid keep / rewrite / kill

Rewritten criteria: [the judgment/verification/outcomes/leverage set, with observable definitions each]

Evidence to collect: [the walk-backwards sample protocol + the rest]

Calibration rules: [the mixed-adoption rules, as committee guidance]

The conversations: [scripts for the three hard cases, adapted to the situation given]

Quality Checks

  • Every current criterion has a human/tool/hybrid verdict — none skipped as "obviously fine"
  • New criteria are observable behaviours, not virtues ("catches errors before shipping" not "is diligent")
  • Verification labour is explicitly valued somewhere — the invisible work made visible
  • Calibration rules prevent both punishing adoption and punishing non-adoption
  • The launderer case routes to reliability/accountability, not to relitigating the AI policy

Anti-Patterns

  • Do not credit or blame the human for what the model did — walk the work backwards to find the human
  • Do not keep volume metrics "because they're objective" — they're objective measurements of the wrong thing now
  • Do not run calibration comparing raw output across uneven adopters — that's a tooling lottery, not a review
  • Do not treat AI scepticism as a performance problem where use is optional — outcomes are the bar, not enthusiasm
  • Do not have the accountability conversation without the org's policy in hand — improvised rules in a review are how grievances are born

Version History

  • a38bc30 Current 2026-07-05 11:29

Same Skill Collection

exports/openclaw/360-feedback-template/SKILL.md
exports/openclaw/401k-plan-decoder/SKILL.md
exports/openclaw/ab-test-planner/SKILL.md
exports/openclaw/ab-test-readout/SKILL.md
exports/openclaw/accessibility-audit/SKILL.md
exports/openclaw/account-plan/SKILL.md
exports/openclaw/acquirer-red-team/SKILL.md
exports/openclaw/ad-copy/SKILL.md
exports/openclaw/aeo-optimizer/SKILL.md
exports/openclaw/agenda-or-cancel/SKILL.md
exports/openclaw/agent-design-review/SKILL.md
exports/openclaw/agent-observability-spec/SKILL.md
exports/openclaw/agent-spec/SKILL.md
exports/openclaw/ai-ethics-review/SKILL.md
exports/openclaw/ai-eval-plan/SKILL.md
exports/openclaw/ai-feature-prd/SKILL.md
exports/openclaw/ai-product-canvas/SKILL.md
exports/openclaw/air-quality/SKILL.md
exports/openclaw/altitude-shifter/SKILL.md
exports/openclaw/ambiguity-resolver/SKILL.md
exports/openclaw/analyst-relations-brief/SKILL.md
exports/openclaw/announcement-card/SKILL.md
exports/openclaw/api-docs-writer/SKILL.md
exports/openclaw/api-test-plan/SKILL.md
exports/openclaw/api-versioning-strategy/SKILL.md
exports/openclaw/apology-letter/SKILL.md
exports/openclaw/architecture-decision-record/SKILL.md
exports/openclaw/architecture-diagram/SKILL.md
exports/openclaw/archive-strategy/SKILL.md
exports/openclaw/assumption-bounty/SKILL.md
exports/openclaw/assumption-mapper/SKILL.md
exports/openclaw/async-update-format/SKILL.md
exports/openclaw/auto-repair-estimate-decoder/SKILL.md
exports/openclaw/autopilot-charter/SKILL.md
exports/openclaw/benefits-decoder/SKILL.md
exports/openclaw/bid-tender-review/SKILL.md
exports/openclaw/board-deck-narrative/SKILL.md
exports/openclaw/board-minutes/SKILL.md
exports/openclaw/board-pre-read/SKILL.md
exports/openclaw/bom-cost-review/SKILL.md
exports/openclaw/bookkeeping-categorization/SKILL.md
exports/openclaw/boolean-search-builder/SKILL.md
exports/openclaw/brag-doc/SKILL.md
exports/openclaw/brainstorming/SKILL.md
exports/openclaw/brief-builder/SKILL.md
exports/openclaw/briefing-note/SKILL.md
exports/openclaw/budget-builder/SKILL.md
exports/openclaw/budget-variance-analysis/SKILL.md
exports/openclaw/bug-diagnosis/SKILL.md
exports/openclaw/bug-report/SKILL.md

Metadata

Files
0
Version
471c606
Hash
cdb460f5
Indexed
2026-07-05 11:29

Accueil - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-07-30 07:00
浙ICP备14020137号-1 $Carte des visiteurs$