Agent Skills
› firecrawl/cli
› firecrawl-crawl
firecrawl-crawl
GitHub用于批量抓取网站或特定目录下的所有页面内容,支持深度限制、路径过滤和并发提取。适用于需要获取大量同一站点内容的场景,如文档库爬取。
Trigger Scenarios
用户要求爬取整个网站或特定部分
用户提及 'crawl', 'get all the pages', 'bulk extract' 等关键词
需要从多个页面提取内容
Install
npx skills add firecrawl/cli --skill firecrawl-crawl -g -y
SKILL.md
Frontmatter
{
"name": "firecrawl-crawl",
"description": "Bulk extract content from an entire website or site section. Use this skill when the user wants to crawl a site, extract all pages from a docs section, bulk-scrape multiple pages following links, or says \"crawl\", \"get all the pages\", \"extract everything under \/docs\", \"bulk extract\", or needs content from many pages on the same site. Handles depth limits, path filtering, and concurrent extraction.\n",
"allowed-tools": [
"Bash(firecrawl *)",
"Bash(npx firecrawl-cli *)"
]
}
firecrawl crawl
Bulk extract content from a website. Crawls pages following links up to a depth/limit.
Prerequisite: crawl requires authentication (no keyless free tier); without credentials the CLI prompts an interactive login.
When to use
- You need content from many pages on a site (e.g., all
/docs/) - You want to extract an entire site section
- Step 4 in the workflow escalation pattern: search → scrape → map + scrape → crawl → monitor → interact
Quick start
# Crawl a docs section
firecrawl crawl "<url>" --include-paths /docs --limit 50 --wait -o .firecrawl/crawl.json
# Full crawl with depth limit
firecrawl crawl "<url>" --max-depth 3 --wait --progress -o .firecrawl/crawl.json
# Check status of a running crawl
firecrawl crawl <job-id>
Options
| Option | Description |
|---|---|
--wait |
Wait for crawl to complete before returning |
--progress |
Show progress while waiting |
--limit <n> |
Max pages to crawl |
--max-depth <n> |
Max link depth to follow |
--include-paths <paths> |
Only crawl URLs matching these paths |
--exclude-paths <paths> |
Skip URLs matching these paths |
--delay <ms> |
Delay between requests |
--max-concurrency <n> |
Max parallel crawl workers |
--pretty |
Pretty print JSON output |
-o, --output <path> |
Output file path |
Tips
- Always use
--waitwhen you need the results immediately. It has no default timeout; use--timeout <seconds>to bound polling. Without--wait, crawl returns a job ID for async polling. - Use
--include-pathsto scope the crawl — don't crawl an entire site when you only need one section. - Crawl consumes credits per page. Check
firecrawl credit-usagebefore large crawls (credit-usagerequires authentication).
See also
- firecrawl-scrape — scrape individual pages
- firecrawl-map — discover URLs before deciding to crawl
- firecrawl-download — download site to local files (uses map + scrape)
Version History
-
253abde
Current 2026-08-19 19:06
新增认证前置说明;更新工作流步骤描述;修复链接并优化文档结构。
- 6c50c5d 2026-07-24 11:49


