Agent Skillspaperclipai/paperclip › paperclip-evals

paperclip-evals

GitHub

用于选择、检查、验证 Paperclip Runner 或产品 E2E 评估,确保保留证据链、溯源信息、成本及失败分类。

.agents/skills/paperclip-evals/SKILL.md paperclipai/paperclip

Trigger Scenarios

需要执行或查看 Paperclip 评估结果 分析 E2E 测试失败原因

Install

npx skills add paperclipai/paperclip --skill paperclip-evals -g -y
More Options

Non-standard path

npx skills add https://github.com/paperclipai/paperclip/tree/master/.agents/skills/paperclip-evals -g -y

Use without installing

npx skills use paperclipai/paperclip@paperclip-evals

指定 Agent (Claude Code)

npx skills add paperclipai/paperclip --skill paperclip-evals -a claude-code -g -y

安装 repo 全部 skill

npx skills add paperclipai/paperclip --all -g -y

预览 repo 内 skill

npx skills add paperclipai/paperclip --list

SKILL.md

Frontmatter
{
    "name": "paperclip-evals",
    "description": "Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification."
}

Paperclip evals

Use this skill when a request concerns Paperclip evaluation selection, interpretation, evidence, history, or a live run. Read doc/evals.md in the Paperclip repository first. It defines the two families and their boundaries.

Discover the repository

Do not assume the skill's installed location is inside a checkout. Locate the repo explicitly with git rev-parse --show-toplevel from the current directory, or inspect likely workspace roots and select the checkout containing package.json, tests/runner-e2e, and packages/paperclip-runner. Locate the private sibling paperclip-evals only when a Runner Eval needs its definitions; use an explicit PAPERCLIP_EVALS_ROOT or a discovered sibling checkout. Never invent a relative path from this copied skill into the repository.

Route the request

Choose Runner Evals for real runner/provider protocol behavior against the mock control plane. Authoritative details are in packages/paperclip-runner/docs/runner-protocol-live-evals.md and the sibling paperclip-evals/evals/paperclip-runner definitions.

Choose Product E2E Evals for real browser/server/database/runner/provider workflows, including local and Daytona environments. Read tests/runner-e2e/README.md, then FIXTURES.md, SECURITY.md, or EVERYDAY-WORKFLOWS.md as relevant. Everyday Workflows remain Product E2E even when imported into Evalbook. “Headless” is a browser mode, not a family.

Work safely

Start with read-only catalog inspection and credential-free validation. For Product E2E use pnpm test:e2e:runner:typecheck, pnpm test:e2e:runner:unit, and pnpm test:e2e:runner -- --list; run one explicit cell only when the user has authorized a live/paid run and the needed credentials and immutable Daytona image are configured. For Runner Evals use the pinned eval revision and the documented workflow/CLI. Never use a partial selector as evidence of full coverage.

Keep source revisions, definition/catalog fingerprints, model/profile, environment, selected cells, retries, timing, usage/cost coverage, and grader version attached to every interpretation. Preserve partial attempts and classify failures as product, model/provider behavior, grading/evidence, or infrastructure from the observed failure and supported cause. A usable completed behavior failure is not infrastructure; missing provider/profile, transport, startup, or evidence requires examining the evidence before choosing the cause.

Use the existing family generator and viewer. Public projections may contain sanitized fixture conversation and allowlisted tool outcomes/evidence; follow the family's projection and publisher checks. Do not expose raw trusted artifacts, credentials, secrets, private data, provider session IDs, or hidden reasoning. A refresh from retained evidence has zero provider calls and remains the original measurement with a new presentation. Link the public histories and hub from doc/evals.md when reporting results.

For adding a case or fixture, use the narrower add-runner-eval or add-product-e2e-eval skill.

Version History

  • 8326e33 Current 2026-09-22 20:26

Same Skill Collection

.agents/skills/add-product-e2e-eval/SKILL.md
.agents/skills/add-runner-eval/SKILL.md
.agents/skills/check-pr/SKILL.md
.agents/skills/company-creator/SKILL.md
.agents/skills/create-agent-adapter/SKILL.md
.agents/skills/create-issue-interaction-ui/SKILL.md
.agents/skills/create-paperclip-bundled-skill/SKILL.md
.agents/skills/deal-with-security-advisory/SKILL.md
.agents/skills/diagnose-why-work-stopped/SKILL.md
.agents/skills/doc-maintenance/SKILL.md
.agents/skills/garden-inbox/SKILL.md
.agents/skills/paperclip-create-plugin/SKILL.md
.agents/skills/paperclip-dev-workspace-run-verify-fix/SKILL.md
.agents/skills/paperclip-page/SKILL.md
.agents/skills/pr-gardening/SKILL.md
.agents/skills/pr-report/SKILL.md
.agents/skills/prcheckloop/SKILL.md
.agents/skills/prepare-paperclip-pr/SKILL.md
.agents/skills/release-changelog-discord-message/SKILL.md
.agents/skills/release-changelog/SKILL.md
.agents/skills/release/SKILL.md
.agents/skills/terminal-bench-loop/SKILL.md
packages/adapters/hermes/skills/paperclip-task-bridge/SKILL.md
packages/plugins/plugin-llm-wiki/skills/index-refresh/SKILL.md
packages/plugins/plugin-llm-wiki/skills/paperclip-distill/SKILL.md
packages/plugins/plugin-llm-wiki/skills/wiki-ingest/SKILL.md
packages/plugins/plugin-llm-wiki/skills/wiki-lint/SKILL.md
packages/plugins/plugin-llm-wiki/skills/wiki-maintainer/SKILL.md
packages/plugins/plugin-llm-wiki/skills/wiki-query/SKILL.md
skills-releases/paperclip/v0/SKILL.md
skills-releases/paperclip/v7-roster/SKILL.md
skills/agentmail/SKILL.md
skills/paperclip-board/SKILL.md
skills/paperclip-converting-plans-to-tasks/SKILL.md
skills/paperclip-create-agent/SKILL.md
skills/paperclip/SKILL.md
skills/para-memory-files/SKILL.md
.claude/skills/design-guide/SKILL.md

Metadata

Files
0
Version
8326e33
Hash
0b220dd2
Indexed
2026-09-22 20:26

Главная - Вики-сайт
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-22 20:47
浙ICP备14020137号-1