Agent Skillsxberg-io/xberg › batch-extraction

batch-extraction

GitHub

批量提取技能,支持对多文件并行处理、共享配置及单文件覆盖,具备错误恢复机制。适用于目录或通配符文档的一次性结构化提取,可控制并发数与输出布局。

plugin/skills/batch-extraction/SKILL.md xberg-io/xberg

Trigger Scenarios

需要同时从多个文件中提取内容 执行批量文档解析任务 处理带有不同配置的混合格式文件集合

Install

npx skills add xberg-io/xberg --skill batch-extraction -g -y
More Options

Non-standard path

npx skills add https://github.com/xberg-io/xberg/tree/main/plugin/skills/batch-extraction -g -y

Use without installing

npx skills use xberg-io/xberg@batch-extraction

指定 Agent (Claude Code)

npx skills add xberg-io/xberg --skill batch-extraction -a claude-code -g -y

安装 repo 全部 skill

npx skills add xberg-io/xberg --all -g -y

预览 repo 内 skill

npx skills add xberg-io/xberg --list

SKILL.md

Frontmatter
{
    "name": "batch-extraction",
    "description": "Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command, `--file-configs`, `--max-concurrent`, and output layout."
}

Batch extraction

Use this when processing a directory or glob of documents in one pass. xberg batch shares one extraction config across every file, runs extractions concurrently, and returns one structured array — failures on individual files do not abort the run.

Basic usage

# Glob expands to many paths; results come back as a JSON array (default)
xberg batch *.pdf

# Mixed formats, markdown content for LLM ingestion
xberg batch docs/*.docx --content-format markdown

# Recurse with the shell, then extract
xberg batch $(find ./corpus -name '*.pdf')

batch defaults to --format json (vs --format text for single extract). Each array entry is a full extraction result, so downstream code can index by position into the input path list.

xberg batch reports/*.pdf \
  | jq '.[] | {chars: (.content | length), mime: .mime_type}'

Parallelism

--max-concurrent caps how many files extract at once. When omitted, the scheduler derives document concurrency from the total thread budget. Lower it on memory-constrained hosts or when OCR/ML models are active, since each in-flight extraction holds its own buffers. Layout-heavy batches are further limited (1 concurrent extraction for all-PDF-layout batches, 2 for mixed layout):

# Cap at 4 concurrent extractions
xberg batch scans/*.pdf --ocr true --max-concurrent 4

--max-threads additionally caps total internal threads (Rayon, ONNX intra-op, the batch semaphore) for tightly constrained environments:

xberg batch *.pdf --max-concurrent 2 --max-threads 4

Per-file config overrides

A single shared config does not always fit. --file-configs points at a JSON file mapping each path to its own override object, merged on top of the shared config for that file only:

{
  "scan.pdf": { "force_ocr": true },
  "report.pdf": { "output_format": "markdown" },
  "data.xlsx": { "output_format": "json" }
}
xberg batch scan.pdf report.pdf data.xlsx --file-configs overrides.json

Keys are file paths (matching the paths passed on the command line); values are per-file extraction config objects in snake_case, the same shape as a config file.

Output layout

For text/toon output with image extraction, --output-dir controls where referenced image files (e.g. image_0.png) are written; the directory must already exist. JSON output embeds image bytes inline and ignores --output-dir.

mkdir -p out/images
xberg batch slides/*.pptx --extract-images true --output-dir out/images --format text

Error recovery

Batch extraction is fault-tolerant per file: one unreadable or corrupt document does not stop the rest. Inspect results for partial content and surfaced errors rather than relying on the process exit code alone. Pair with --max-concurrent to avoid exhausting memory when a few large files sit in a big batch.

Shared config

Every extract flag also applies to batch (OCR, chunking, layout, content format, etc.) and is shared across all files unless a --file-configs entry overrides it:

xberg batch invoices/*.pdf \
  --layout --layout-table-model slanet_wireless \
  --content-format markdown --max-concurrent 8

A config file works too and auto-discovers from the cwd upward:

output_format = "markdown"

[ocr]
backend = "tesseract"
language = "eng"
xberg batch corpus/*.pdf --config xberg.toml

Programmatic access

From Python, extract_batch takes a list of ExtractInputs and returns one envelope whose results array holds a document per input:

from xberg import ExtractInput, extract_batch, ExtractionConfig

config = ExtractionConfig(output_format="markdown")

inputs = [ExtractInput(uri=p) for p in ["a.pdf", "b.docx", "c.xlsx"]]
output = await extract_batch(inputs, config)

for doc in output.results:
    print(len(doc.content))

Per-input overrides go on ExtractInput.config (a FileExtractionConfig). Node.js mirrors this with extractBatch; Rust uses extract_batch(inputs, &config). See references/python-api.md, references/nodejs-api.md, and references/rust-api.md in the sibling xberg skill.

MCP

When the xberg MCP server is registered, prefer the extract_batch tool over shelling out — it takes an array of input objects and a config object and returns structured results directly.

Common pitfalls

  • Default format differsbatch defaults to --format json, extract to --format text. Set --format explicitly if a script depends on one shape.
  • --output-dir must exist — the CLI does not create it.
  • Memory blowups — large batches with OCR/layout active may need an explicit --max-concurrent ceiling.
  • --file-configs path keys — must match the paths as passed on the command line, not absolute-resolved variants.

See references/cli-reference.md for the full batch flag set.

Version History

  • d8e4815 Current 2026-08-28 18:32

    更新并发调度默认行为描述;完善错误恢复章节说明。

  • 531e0f7 2026-08-20 07:48

Same Skill Collection

.ai-rulez/skills/alef-generated-bindings/SKILL.md
.ai-rulez/skills/benchmark-workflow/SKILL.md
.ai-rulez/skills/chunking-embeddings/SKILL.md
.ai-rulez/skills/config-loading-precedence/SKILL.md
.ai-rulez/skills/crate-structure/SKILL.md
.ai-rulez/skills/extraction-pipeline-patterns/SKILL.md
.ai-rulez/skills/feature-flag-policy/SKILL.md
.ai-rulez/skills/mime-detection-routing/SKILL.md
.ai-rulez/skills/ocr-pipeline-and-quality/SKILL.md
.ai-rulez/skills/pdf-backends/SKILL.md
.ai-rulez/skills/plugin-architecture-patterns/SKILL.md
.ai-rulez/skills/polyrepo-boundaries/SKILL.md
.ai-rulez/skills/release-readiness/SKILL.md
.ai-rulez/skills/release-versioning/SKILL.md
.ai-rulez/skills/test-corpus/SKILL.md
.ai-rulez/skills/wasm-constraints/SKILL.md
.ai-rulez/skills/xberg-typescript-toolchain/SKILL.md
plugin/.ai-rulez/skills/batch-extraction/SKILL.md
plugin/.ai-rulez/skills/chunking/SKILL.md
plugin/.ai-rulez/skills/extracting-keywords/SKILL.md
plugin/.ai-rulez/skills/extracting-tables/SKILL.md
plugin/.ai-rulez/skills/extracting-with-ocr/SKILL.md
plugin/.ai-rulez/skills/picking-a-format/SKILL.md
plugin/.ai-rulez/skills/xberg/SKILL.md
plugin/.cursor-plugin/skills/batch-extraction/SKILL.md
plugin/.cursor-plugin/skills/chunking/SKILL.md
plugin/.cursor-plugin/skills/extracting-keywords/SKILL.md
plugin/.cursor-plugin/skills/extracting-tables/SKILL.md
plugin/.cursor-plugin/skills/extracting-with-ocr/SKILL.md
plugin/.cursor-plugin/skills/picking-a-format/SKILL.md
plugin/.cursor-plugin/skills/xberg/SKILL.md
plugin/skills/chunking/SKILL.md
plugin/skills/extracting-keywords/SKILL.md
plugin/skills/extracting-tables/SKILL.md
plugin/skills/extracting-with-ocr/SKILL.md
plugin/skills/picking-a-format/SKILL.md
plugin/skills/xberg/SKILL.md
.ai-rulez/skills/api-server-mcp/SKILL.md
.ai-rulez/skills/format-specific-extraction/SKILL.md

Metadata

Files
0
Version
d8e4815
Hash
1fb29533
Indexed
2026-08-20 07:48

trang chủ - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-01 02:08
浙ICP备14020137号-1 $bản đồ khách truy cập$