Agent Skills
› anthropics/cwc-workshops
› incident-triage-runbook
incident-triage-runbook
GitHubSRE生产事故排查手册,指导处理延迟和错误率异常。按顺序检查部署、指标关联、代码变更及日志,识别常见故障模式并生成根因摘要,辅助快速定位问题。
Trigger Scenarios
调查线上事故
分析延迟飙升
错误率升高
询问服务异常原因
Install
npx skills add anthropics/cwc-workshops --skill incident-triage-runbook -g -y
SKILL.md
Frontmatter
{
"name": "incident-triage-runbook",
"description": "The SRE team's runbook for triaging production latency and error-rate incidents. Use this whenever investigating an incident, a latency spike, elevated error rates, or when asked \"what caused X\" about a production service."
}
Incident triage
If you change the order below, say why in #sre.
Order of operations
- Pull deploys for the last 6h. Don't open the log first.
- Line the deploy timestamps up against
p99_latency_ms/error_ratefor the paged service. State the gap ("deploy 14:31, p99 moves 14:33"). - If a deploy lines up: pull the diff, read it. Check for the stuff in the next section.
- Then grep the log to confirm. Don't grep to fish.
- No deploy lines up → check
db_pool_utilizationacross checkout/cart/auth/inventory, then upstream deps.
Things that have burned us
In rough order of how often:
- per-row query where there used to be a batch
- cache decorator removed "temporarily"
- new query, no index
- blocking call in an async handler
- retry loop with no backoff
Write-up
One line at the bottom:
Root cause:
<sha>— one sentence on the mechanism.
If it wasn't a deploy, put the component or upstream dep where the sha goes (db-primary, stripe-api, whatever). Still one sentence.
Everything above that line is evidence. Keep it short; the long version goes in the postmortem doc.
Version History
- 068b84b Current 2026-08-27 17:23


