Agent Skillsxberg-io/xberg › benchmark-workflow

benchmark-workflow

GitHub

管理基准测试工作流,涵盖运行、诊断及变更提取基准、质量评分和固定装置。强调使用独立源验证真相完整性,区分基础设施与提取故障,确保A/B测试环境一致性。

.ai-rulez/skills/benchmark-workflow/SKILL.md xberg-io/xberg

Trigger Scenarios

需要运行或分析基准测试结果时 调整基准测试配置或修复质量评分问题时 排查提取器或OCR引擎的质量缺陷时

Install

npx skills add xberg-io/xberg --skill benchmark-workflow -g -y
More Options

Non-standard path

npx skills add https://github.com/xberg-io/xberg/tree/main/.ai-rulez/skills/benchmark-workflow -g -y

Use without installing

npx skills use xberg-io/xberg@benchmark-workflow

指定 Agent (Claude Code)

npx skills add xberg-io/xberg --skill benchmark-workflow -a claude-code -g -y

安装 repo 全部 skill

npx skills add xberg-io/xberg --all -g -y

预览 repo 内 skill

npx skills add xberg-io/xberg --list

SKILL.md

Frontmatter
{
    "name": "benchmark-workflow",
    "description": "Run, diagnose, or change Xberg extraction benchmarks, quality scoring, benchmark fixtures, artifact contracts, and independently sourced ground truth. Load for the Benchmarks workflow or benchmark-harness work, not ordinary unit tests."
}

Benchmark workflow

The benchmark system lives in tools/benchmark-harness/; the GitHub workflow is .github/workflows/benchmarks.yaml. The workflow is dispatch-only, so it does not run on push or gate merges. Treat a result as evidence for its exact commit SHA and inputs, not for newer local work.

Ground-truth integrity

  • Never use Xberg's own extractor output as benchmark ground truth. Use an independent source and record it in the fixture's ground_truth.source field (manual, vision, pdf_text_layer, pandoc, python-docx, and similar).
  • Before blaming ground truth for a score, render or otherwise inspect the source document. If the derived .md or .txt disagrees with the source, fix the ground truth; if it agrees, investigate the extractor or metric.
  • Use the fixture schema in tools/benchmark-harness/README.md, the generator at tools/benchmark-harness/scripts/generate_markdown_gt.py, and the harness validate-gt command implemented in tools/benchmark-harness/src/validate_gt.rs. Do not replace these with an ad-hoc conversion pipeline.
  • A quality claim requires the same corpus, config, renderer, cache state, and metric on control and experiment. Disable or invalidate extraction and OCR caches before A/B runs whose output behavior changed.

Diagnosing runs

  • Separate infrastructure failures from extraction or quality failures. A missing backend library, absent fixture, malformed artifact, or runner setup error does not describe extractor quality.
  • Inspect the per-adapter artifacts before the aggregate job. Aggregate contract failures may be caused by missing or unexpectedly named artifacts even when individual adapters ran.
  • Compare accepted OCR pages before raw word counts. Rejected OCR pages contribute neither text nor structured paragraphs.
  • Measure headings and lists using Markdown output. Plain output normalizes away list markers and cannot distinguish detection from rendering.
  • Do not quote a coverage, latency, or quality threshold unless the workflow or harness currently enforces it.

When changing harness behavior, add focused tests for the report or artifact contract and run the task that exercises the affected adapter before dispatching the remote workflow.

Version History

  • d8e4815 Current 2026-08-28 18:30

Same Skill Collection

.ai-rulez/skills/alef-generated-bindings/SKILL.md
.ai-rulez/skills/chunking-embeddings/SKILL.md
.ai-rulez/skills/config-loading-precedence/SKILL.md
.ai-rulez/skills/crate-structure/SKILL.md
.ai-rulez/skills/extraction-pipeline-patterns/SKILL.md
.ai-rulez/skills/feature-flag-policy/SKILL.md
.ai-rulez/skills/mime-detection-routing/SKILL.md
.ai-rulez/skills/ocr-pipeline-and-quality/SKILL.md
.ai-rulez/skills/pdf-backends/SKILL.md
.ai-rulez/skills/plugin-architecture-patterns/SKILL.md
.ai-rulez/skills/polyrepo-boundaries/SKILL.md
.ai-rulez/skills/release-readiness/SKILL.md
.ai-rulez/skills/release-versioning/SKILL.md
.ai-rulez/skills/test-corpus/SKILL.md
.ai-rulez/skills/wasm-constraints/SKILL.md
.ai-rulez/skills/xberg-typescript-toolchain/SKILL.md
plugin/.ai-rulez/skills/batch-extraction/SKILL.md
plugin/.ai-rulez/skills/chunking/SKILL.md
plugin/.ai-rulez/skills/extracting-keywords/SKILL.md
plugin/.ai-rulez/skills/extracting-tables/SKILL.md
plugin/.ai-rulez/skills/extracting-with-ocr/SKILL.md
plugin/.ai-rulez/skills/picking-a-format/SKILL.md
plugin/.ai-rulez/skills/xberg/SKILL.md
plugin/.cursor-plugin/skills/batch-extraction/SKILL.md
plugin/.cursor-plugin/skills/chunking/SKILL.md
plugin/.cursor-plugin/skills/extracting-keywords/SKILL.md
plugin/.cursor-plugin/skills/extracting-tables/SKILL.md
plugin/.cursor-plugin/skills/extracting-with-ocr/SKILL.md
plugin/.cursor-plugin/skills/picking-a-format/SKILL.md
plugin/.cursor-plugin/skills/xberg/SKILL.md
plugin/skills/batch-extraction/SKILL.md
plugin/skills/chunking/SKILL.md
plugin/skills/extracting-keywords/SKILL.md
plugin/skills/extracting-tables/SKILL.md
plugin/skills/extracting-with-ocr/SKILL.md
plugin/skills/picking-a-format/SKILL.md
plugin/skills/xberg/SKILL.md
.ai-rulez/skills/api-server-mcp/SKILL.md
.ai-rulez/skills/format-specific-extraction/SKILL.md

Metadata

Files
0
Version
971d738
Hash
6802ba9f
Indexed
2026-08-28 18:30

inicio - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-11 01:43
浙ICP备14020137号-1 $mapa de visitantes$