proofloop
GitHub通过执行行为检查、可重复工件及独立GPT-6 Luna模型验证代码变更,合并低信号测试。提供从设置检查器到定义检查策略再到执行候选方案的完整流程,实现基于证据的测试精简与质量保障。
Trigger Scenarios
Install
npx skills add regenrek/codex-proofloop --skill proofloop -g -y
SKILL.md
Frontmatter
{
"name": "proofloop",
"description": "Verify coding changes with executed behavioral checks, repeatable artifacts and an independent native GPT-6 Luna checker at max reasoning. Consolidate low-signal tests. Use when the user requests Proofloop or evidence-based test reduction."
}
Proofloop
Prefer a small set of repeatable checks of actual behavior. Temporary diagnostic probes do not
need to become permanent tests. The current task owns implementation; a native gpt-6-luna task with
max reasoning independently checks the candidate by default. Keep the current implementation model.
The user supplies the project outcome; this skill supplies the checker model and coordination.
Keep verification proportional to the change. Spend review effort on concrete failure modes the selected checks could miss. No mandatory audit report, feedback diary or test-deletion quota.
Set up the checker
Read orchestration.md before starting a run. Reuse a suitable existing project checker task; when task creation is authorized, create the missing checker using the native host tools. Apply standing user authorization without asking again. Do not ask the user to repeat the default model, reasoning level or checker role in each project prompt.
Resolve the actual checker task ID and return task ID before start, then fill the policy's review
assignment. The empty ID in the template deliberately requires setup; never invent an ID. If host
capabilities or required task-creation authorization are missing, report that specific blocker.
The skill does not override host authorization rules. Never silently drop the checker or substitute
another model. Use review: null only when the user explicitly chooses execution without independent
review, and identify that limitation in the handoff.
Define the check before editing
Read applicable project instructions and existing tests. State the intended user-visible behavior,
concrete failure modes and the checks that cover them in one proofloop.json, using
the policy template. Select commands the project actually provides.
Do not infer a full-suite command or invent an available environment. Define the target and synthetic
seed; use no credentials in policy or artifacts. Add .proofloop/ to the project's root .gitignore.
Prefer E2E at the public boundary: browser for a web journey, real processes for a CLI. A permanent isolated test needs a concrete failure that the selected E2E cannot adequately expose. List those failures before implementing it. Do not append unit tests that restate newly written code.
For existing tests, read testing.md. Declare each intended add, modification
or deletion in testChanges with its reason and surviving checks. Unknown coverage is not proof of
redundancy. Keep unrelated user changes.
Execute against one candidate
Use the absolute path to this installed skill's scripts/cli.mjs below. The skill includes compiled
JavaScript and needs only Node 24+ and Git; installation does not start anything.
node /absolute/skill/path/scripts/cli.mjs start --project /absolute/project --id task-123
# Implement within the frozen policy. Temporary probes may live under .proofloop/.
node /absolute/skill/path/scripts/cli.mjs run --project /absolute/project --id task-123 --compact
# After the assigned checker returns its actual review:
node /absolute/skill/path/scripts/cli.mjs finish --project /absolute/project --id task-123 --review .proofloop/checker-response.json --compact
Read runner.md when selecting reporters, artifact paths or diagnosing an incomplete run. Never edit generated run records to make a check pass. A policy change requires a new run ID. Source changes after a check require rerunning the checks. An intermediate commit does not hide file changes from the baseline.
Use native vitest-json for existing Vitest suites and exit-code for build/lint/typecheck commands;
do not write a Node-test wrapper just to make these commands fit. Each criterion still needs a
behavioral check. Direct local file arguments are hashed even when ignored. Declare additional
ignored helpers, configs and fixtures in each check's inputs; the runner does not trace imports
or package scripts. Keep outputs out of inputs.
The runner records process results; it cannot decide whether an assertion represents the right product requirement. Classify failures as a real defect, wrong expectation or low-value check before changing production behavior. Do not relax the policy just to obtain a green result.
Independent check before completion
After execution, hand the frozen candidate and evidence to the assigned Luna checker using
orchestration.md. Collect its actual response before completing the
run with finish --review. A missing required checker leaves the run incomplete. Do not add Herdr,
a watcher, or another manager.
Completion
Use status --compact for freshness and hand off its record path, relevant diff and a short objective.
The runner stores commands, hashes and artifacts; do not transcribe them into a second report. Read
full records and specific log sections only as needed. Run the selected checks once per candidate;
avoid a duplicate preflight of the same commands.
Return a short outcome, actionable findings or important limits, and the final record path. Mention
test additions/deletions with their reason when applicable. Stop once checks and the independent
review pass. Do not add another audit or generate a feedback report unless requested. status
recalculates freshness; a historical finish.json does not establish freshness after further edits.
Version History
- 7c81dab Current 2026-09-27 12:02


