Agent Skillsxberg-io/xberg › benchmark-workflow

benchmark-workflow

GitHub

管理基准测试工作流,涵盖提取基准、质量评分及断言契约。指导如何维护独立源真值、诊断运行故障(区分基础设施与提取问题)及验证结果,确保评估准确性。

.ai-rulez/skills/benchmark-workflow/SKILL.md xberg-io/xberg

Trigger Scenarios

需要运行或诊断基准测试时 修改基准夹具或质量评分逻辑时 排查提取器或OCR质量问题时

Install

npx skills add xberg-io/xberg --skill benchmark-workflow -g -y
More Options

Non-standard path

npx skills add https://github.com/xberg-io/xberg/tree/main/.ai-rulez/skills/benchmark-workflow -g -y

Use without installing

npx skills use xberg-io/xberg@benchmark-workflow

指定 Agent (Claude Code)

npx skills add xberg-io/xberg --skill benchmark-workflow -a claude-code -g -y

安装 repo 全部 skill

npx skills add xberg-io/xberg --all -g -y

预览 repo 内 skill

npx skills add xberg-io/xberg --list

SKILL.md

Frontmatter
{
    "name": "benchmark-workflow",
    "description": "Run, diagnose, or change Xberg extraction benchmarks, quality scoring, benchmark fixtures, artifact contracts, and independently sourced ground truth. Load for the Benchmarks workflow or benchmark-harness work, not ordinary unit tests."
}

Benchmark workflow

The benchmark system lives in tools/benchmark-harness/; the GitHub workflow is .github/workflows/benchmarks.yaml. The workflow is dispatch-only, so it does not run on push or gate merges. Treat a result as evidence for its exact commit SHA and inputs, not for newer local work.

Ground-truth integrity

  • Never use Xberg's own extractor output as benchmark ground truth. Use an independent source and record it in the fixture's ground_truth.source field (manual, vision, pdf_text_layer, pandoc, python-docx, and similar).
  • Before blaming ground truth for a score, render or otherwise inspect the source document. If the derived .md or .txt disagrees with the source, fix the ground truth; if it agrees, investigate the extractor or metric.
  • Use the fixture schema in tools/benchmark-harness/README.md, the generator at tools/benchmark-harness/scripts/generate_markdown_gt.py, and the harness validate-gt command implemented in tools/benchmark-harness/src/validate_gt.rs. Do not replace these with an ad-hoc conversion pipeline.
  • A quality claim requires the same corpus, config, renderer, cache state, and metric on control and experiment. Disable or invalidate extraction and OCR caches before A/B runs whose output behavior changed.

Diagnosing runs

  • Separate infrastructure failures from extraction or quality failures. A missing backend library, absent fixture, malformed artifact, or runner setup error does not describe extractor quality.
  • Inspect the per-adapter artifacts before the aggregate job. Aggregate contract failures may be caused by missing or unexpectedly named artifacts even when individual adapters ran.
  • Compare accepted OCR pages before raw word counts. Rejected OCR pages contribute neither text nor structured paragraphs.
  • Measure headings and lists using Markdown output. Plain output normalizes away list markers and cannot distinguish detection from rendering.
  • Do not quote a coverage, latency, or quality threshold unless the workflow or harness currently enforces it.

When changing harness behavior, add focused tests for the report or artifact contract and run the task that exercises the affected adapter before dispatching the remote workflow.

Version History

  • d8e4815 Current 2026-08-28 18:30

Same Skill Collection

.ai-rulez/skills/alef-generated-bindings/SKILL.md
.ai-rulez/skills/chunking-embeddings/SKILL.md
.ai-rulez/skills/config-loading-precedence/SKILL.md
.ai-rulez/skills/crate-structure/SKILL.md
.ai-rulez/skills/extraction-pipeline-patterns/SKILL.md
.ai-rulez/skills/feature-flag-policy/SKILL.md
.ai-rulez/skills/mime-detection-routing/SKILL.md
.ai-rulez/skills/ocr-pipeline-and-quality/SKILL.md
.ai-rulez/skills/pdf-backends/SKILL.md
.ai-rulez/skills/plugin-architecture-patterns/SKILL.md
.ai-rulez/skills/polyrepo-boundaries/SKILL.md
.ai-rulez/skills/release-readiness/SKILL.md
.ai-rulez/skills/release-versioning/SKILL.md
.ai-rulez/skills/test-corpus/SKILL.md
.ai-rulez/skills/wasm-constraints/SKILL.md
.ai-rulez/skills/xberg-typescript-toolchain/SKILL.md
plugin/.ai-rulez/skills/batch-extraction/SKILL.md
plugin/.ai-rulez/skills/chunking/SKILL.md
plugin/.ai-rulez/skills/extracting-keywords/SKILL.md
plugin/.ai-rulez/skills/extracting-tables/SKILL.md
plugin/.ai-rulez/skills/extracting-with-ocr/SKILL.md
plugin/.ai-rulez/skills/picking-a-format/SKILL.md
plugin/.ai-rulez/skills/xberg/SKILL.md
plugin/.cursor-plugin/skills/batch-extraction/SKILL.md
plugin/.cursor-plugin/skills/chunking/SKILL.md
plugin/.cursor-plugin/skills/extracting-keywords/SKILL.md
plugin/.cursor-plugin/skills/extracting-tables/SKILL.md
plugin/.cursor-plugin/skills/extracting-with-ocr/SKILL.md
plugin/.cursor-plugin/skills/picking-a-format/SKILL.md
plugin/.cursor-plugin/skills/xberg/SKILL.md
plugin/skills/batch-extraction/SKILL.md
plugin/skills/chunking/SKILL.md
plugin/skills/extracting-keywords/SKILL.md
plugin/skills/extracting-tables/SKILL.md
plugin/skills/extracting-with-ocr/SKILL.md
plugin/skills/picking-a-format/SKILL.md
plugin/skills/xberg/SKILL.md
.ai-rulez/skills/api-server-mcp/SKILL.md
.ai-rulez/skills/format-specific-extraction/SKILL.md

Metadata

Files
0
Version
d8e4815
Hash
6802ba9f
Indexed
2026-08-28 18:30

trang chủ - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-01 21:09
浙ICP备14020137号-1 $bản đồ khách truy cập$