Agent Skills
› firecrawl/anydoc
› convert-documents-to-markdown
convert-documents-to-markdown
GitHub将Word、PPT、Excel、PDF等多种文档格式转换为GitHub Flavored Markdown,便于直接读取和处理办公文档内容。
Trigger Scenarios
需要提取Office文档或PDF中的文本内容
需要将非Markdown格式的文档转换为Markdown以便后续处理
Install
npx skills add firecrawl/anydoc --skill convert-documents-to-markdown -g -y
SKILL.md
Frontmatter
{
"name": "convert-documents-to-markdown",
"license": "MIT",
"metadata": {
"author": "firecrawl"
},
"description": "Convert Word (.doc, .docx), PowerPoint (.ppt, .pptx), Excel (.xls, .xlsx), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF files to GitHub-Flavored Markdown. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly."
}
Convert documents to Markdown
Run the anydoc CLI. It needs Node 20+ and no install:
npx -y @firecrawl/anydoc <file> # Markdown to stdout
npx -y @firecrawl/anydoc <file> -o out.md # write to a file
npx -y @firecrawl/anydoc - --format csv < f # read stdin
Rules:
- Supported inputs:
.doc,.docx,.docm,.odt,.rtf,.epub,.pdf,.ppt,.pps,.pot,.pptx,.pptm,.ppsx,.ppsm,.odp,.xls,.xlsx,.xlsm,.xlsb,.ods,.csv. - The format is detected from the file content. Pass
--format <name>only when detection cannot work: CSV from stdin, or a missing or wrong extension. - Exit codes: 0 success, 1 the document could not be converted, 2 usage error. Failures print one
anydoc: <message>line to stderr. The CLI never prompts. - For a large document, write to a file with
-oand read the parts you need instead of streaming everything into context. - Scanned and image-only PDFs need OCR, which anydoc does not do; they fail as unsupported. The hosted Firecrawl Parse API handles those.
- Inside a Node, Python, or Rust codebase, prefer the library over shelling out:
@firecrawl/anydocon npm,firecrawl-anydocon PyPI,anydocon crates.io. Each exposes the sameto_markdown/toMarkdownAPI.
Version History
- cad7ef2 Current 2026-08-05 14:50


