review-usability
GitHub通过对比 CPython 与 Monty 对常见 Python 代码的执行结果,识别行为差异和静默错误。重点发现影响 LLM 生成代码的兼容性问题和未文档化的分歧,并生成详细报告。
Trigger Scenarios
Install
npx skills add pydantic/monty --skill review-usability -g -y
SKILL.md
Frontmatter
{
"name": "review-usability",
"description": "Check whether the common Python code an LLM would plausibly write still works on this branch, testing real cases in .\/playground against CPython. Use to find behaviour that diverges from CPython or trips up ordinary idiomatic code."
}
Usability review
Monty exists so LLMs can write Python that calls tools. Real usage is therefore the most common patterns, not exotic corners — and a divergence in a common pattern is the worst kind of bug, because the model has no way to know it must write something else.
Think hard on this one.
git diff origin/main...HEAD
-
For each feature the branch touches, list the idioms a model reaches for first — the obvious method, argument form, combination with another builtin. Include ones the branch does not handle; that's where the gaps are.
-
Write real test files in
playground/(seepython-playground), named recognisably. -
Run each under both and diff:
uv run playground/test_thing.py # CPython cargo run -- playground/test_thing.py # Monty -
Prioritise silent divergence — same code, different result — over a clean
AttributeError. A missing feature that raises is recoverable; a wrong answer isn't.
An undocumented divergence is also a ./limitations/ finding.
Report
Per divergence: the code, CPython's output, Monty's output, how likely a model is to write it. Then unsupported-but-common idioms with the error the user sees, and briefly what worked — it bounds the review. Leave the playground files in place.
Report only, unless the user asks for fixes.
Version History
- 026b383 Current 2026-08-20 09:39


