Agent Skills
› cbrock84/headcount
› systematic-debugging
systematic-debugging
GitHub提供系统化调试方法论,强调先复现和定位根因再修复。适用于测试失败、生产错误及间歇性问题排查,指导通过二分法缩小范围、验证假设并确认因果,避免盲目修改。
Trigger Scenarios
测试失败或断言错误
生产环境报错或异常行为
间歇性/随机性问题
修复未生效或问题复发
Install
npx skills add cbrock84/headcount --skill systematic-debugging -g -y
SKILL.md
Frontmatter
{
"name": "systematic-debugging",
"description": "Finds the root cause of a bug, test failure, or unexpected behavior before proposing any fix. Use this whenever something is broken and the cause is not yet proven — a failing test, a production error, intermittent behavior, or a symptom that appeared after a change. Also use when a fix has been attempted and did not work, or when the same bug keeps coming back."
}
Systematic debugging
The rule
No fix before the cause is proven. A change that makes a symptom disappear without an explanation has not fixed anything — it has moved the failure somewhere you are not looking.
Method
- Reproduce it deterministically. If you cannot make it happen on demand, you cannot know when it is fixed. Intermittent means you have not found the variable yet — order, timing, state, environment, data.
- Narrow the blast radius. Bisect: which commit, which input, which branch, which layer. Halve the search space with each step rather than reading everything.
- State a hypothesis that can be wrong. "The cache returns stale rows after a write" is a hypothesis. "Something is wrong with caching" is not.
- Test the hypothesis directly — a log line, a breakpoint, a probe. Prove it, do not infer it.
- Explain the whole symptom. If your cause explains the error but not why it started Tuesday, you have found a bug, not the bug.
- Fix, then verify by reverting. Put the bug back and confirm the test fails again. This is the step people skip, and it is the one that proves causation rather than coincidence.
Anti-patterns
- Shotgun changes — altering several things at once. Now you cannot attribute the fix.
- "Probably a flake." Not a diagnosis. A test that fails intermittently is reporting a real race, ordering dependency, or shared-state leak.
- Fixing the symptom — catching the exception, adding a retry, widening a timeout — without knowing what threw it.
- Trusting the error message's location. Where it surfaced is rarely where it originated.
Never
- Close a bug without a reproduction that failed before the fix and passes after it.
- Deploy a change to production to find out whether it works. Reproduce somewhere you can observe first.
- Change the environment and the code in the same step.
- Leave a debugging aid in the code as the fix — a widened timeout, a disabled check, a swallowed exception.
Return contract
State the reproduction, the proven cause, the fix, and the verification that the fix addresses that cause specifically. Name anything you ruled out and how.
Version History
- d58a7ee Current 2026-09-02 21:11


