Agent Skillsxberg-io/xberg › picking-a-format

picking-a-format

GitHub

指导选择文档提取输出格式(如文本、Markdown、JSON等)及CLI参数组合,适配LLM、RAG或人类审阅等下游消费场景。

plugin/.cursor-plugin/skills/picking-a-format/SKILL.md xberg-io/xberg

Trigger Scenarios

需要决定提取文档的输出格式 配置CLI工具以匹配特定消费者

Install

npx skills add xberg-io/xberg --skill picking-a-format -g -y
More Options

Non-standard path

npx skills add https://github.com/xberg-io/xberg/tree/main/plugin/.cursor-plugin/skills/picking-a-format -g -y

Use without installing

npx skills use xberg-io/xberg@picking-a-format

指定 Agent (Claude Code)

npx skills add xberg-io/xberg --skill picking-a-format -a claude-code -g -y

安装 repo 全部 skill

npx skills add xberg-io/xberg --all -g -y

预览 repo 内 skill

npx skills add xberg-io/xberg --list

SKILL.md

Frontmatter
{
    "name": "picking-a-format",
    "description": "Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` \/ `--content-format` pair."
}

Picking a format

Xberg has two orthogonal format knobs. Get them right up front and the downstream code stays simple.

Knob What it controls Values Default
--format How the CLI prints the result text, json, toon text (extract), json (batch)
--content-format How extracted content is rendered inside result plain, markdown, djot, html, json, doctags plain
--token-reduction Strip whitespace / boilerplate for LLM contexts off, light, moderate, aggressive, maximum off

--format json returns an envelope wrapping the ExtractedDocument — the document lives under .result for extract and under .results[] for batch, with content, metadata, tables, and images as fields of that nested document. --format text prints just content. --content-format is what shows up inside that content field.

Decision tree

Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│       --format text --content-format markdown
├── Vector store / RAG indexer
│       --format json --content-format markdown
│       (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│       --format json --content-format plain
│       (cleanest text + structured metadata)
├── Human review / archival
│       --format text --content-format markdown
├── HTML re-rendering / web display
│       --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│       --format json --content-format djot
└── Token-budget-constrained pipeline
        --format text --content-format plain
        (drops markup; add --token-reduction moderate for further savings)

Examples

Feed a PDF directly into an LLM:

xberg extract paper.pdf --content-format markdown

Index a corpus into a RAG store with tables and headings preserved:

xberg batch docs/*.pdf --format json --content-format markdown \
  | jq -c '.results[] | {content: .content, tables: .tables}'

Strip a file to bare text for a token-tight summarizer:

xberg extract long.pdf \
  --content-format plain \
  --token-reduction moderate

Pull metadata only, ignore content:

xberg extract file.pdf --format json | jq '.result.metadata'

When in doubt

  • Default to markdown as the content format. It is the best compromise across LLMs, RAG, and human review, and Xberg has the most faithful renderer for it.
  • Reach for plain only when downstream cannot tolerate any markup.
  • Reach for djot only if you're already in a djot/pandoc pipeline.
  • Reach for html only when re-rendering for the web.
  • Reach for json for a heading-driven content tree, or doctags for Docling-compatible output.

Token-reduction (orthogonal)

--token-reduction collapses whitespace, strips repeated headers/footers, and trims boilerplate. It composes with any --content-format:

  • off (default), light, moderate, aggressive, maximum.

Use moderate as a safe starting point for LLM context windows. maximum is lossy — verify before relying on it.

See references/cli-reference.md for the full flag set and references/configuration.md for the equivalent output_format and token_reduction keys in xberg.toml.

Version History

  • d8e4815 Current 2026-08-28 18:32

    新增 djot 和 doctags 内容格式选项;更新 token-reduction 说明。

  • 531e0f7 2026-08-20 07:48

Same Skill Collection

.ai-rulez/skills/alef-generated-bindings/SKILL.md
.ai-rulez/skills/benchmark-workflow/SKILL.md
.ai-rulez/skills/chunking-embeddings/SKILL.md
.ai-rulez/skills/config-loading-precedence/SKILL.md
.ai-rulez/skills/crate-structure/SKILL.md
.ai-rulez/skills/extraction-pipeline-patterns/SKILL.md
.ai-rulez/skills/feature-flag-policy/SKILL.md
.ai-rulez/skills/mime-detection-routing/SKILL.md
.ai-rulez/skills/ocr-pipeline-and-quality/SKILL.md
.ai-rulez/skills/pdf-backends/SKILL.md
.ai-rulez/skills/plugin-architecture-patterns/SKILL.md
.ai-rulez/skills/polyrepo-boundaries/SKILL.md
.ai-rulez/skills/release-readiness/SKILL.md
.ai-rulez/skills/release-versioning/SKILL.md
.ai-rulez/skills/test-corpus/SKILL.md
.ai-rulez/skills/wasm-constraints/SKILL.md
.ai-rulez/skills/xberg-typescript-toolchain/SKILL.md
plugin/.ai-rulez/skills/batch-extraction/SKILL.md
plugin/.ai-rulez/skills/chunking/SKILL.md
plugin/.ai-rulez/skills/extracting-keywords/SKILL.md
plugin/.ai-rulez/skills/extracting-tables/SKILL.md
plugin/.ai-rulez/skills/extracting-with-ocr/SKILL.md
plugin/.ai-rulez/skills/picking-a-format/SKILL.md
plugin/.ai-rulez/skills/xberg/SKILL.md
plugin/.cursor-plugin/skills/batch-extraction/SKILL.md
plugin/.cursor-plugin/skills/chunking/SKILL.md
plugin/.cursor-plugin/skills/extracting-keywords/SKILL.md
plugin/.cursor-plugin/skills/extracting-tables/SKILL.md
plugin/.cursor-plugin/skills/extracting-with-ocr/SKILL.md
plugin/.cursor-plugin/skills/xberg/SKILL.md
plugin/skills/batch-extraction/SKILL.md
plugin/skills/chunking/SKILL.md
plugin/skills/extracting-keywords/SKILL.md
plugin/skills/extracting-tables/SKILL.md
plugin/skills/extracting-with-ocr/SKILL.md
plugin/skills/picking-a-format/SKILL.md
plugin/skills/xberg/SKILL.md
.ai-rulez/skills/api-server-mcp/SKILL.md
.ai-rulez/skills/format-specific-extraction/SKILL.md

Metadata

Files
0
Version
d8e4815
Hash
23363f7e
Indexed
2026-08-20 07:48

trang chủ - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-01 02:49
浙ICP备14020137号-1 $bản đồ khách truy cập$