firecrawl-agent
GitHub基于AI自主导航网站并提取结构化数据。优先检查Alexandria预置工作流,无合适工具时启用Agent进行跨页抓取,支持JSON Schema约束输出、异步Job ID管理及信用额度控制。
Trigger Scenarios
Install
npx skills add firecrawl/cli --skill firecrawl-agent -g -y
SKILL.md
Frontmatter
{
"name": "firecrawl-agent",
"description": "Autonomously navigate websites and extract structured data across pages. Use when the task requires navigation or no suitable ready-made workflow or data provider covers it.",
"allowed-tools": [
"Bash(firecrawl *)",
"Bash(npx firecrawl-cli *)"
]
}
firecrawl agent
AI-powered autonomous extraction. The agent navigates sites and extracts structured data (takes 2-5 minutes).
Before starting autonomous extraction for structured records or listings, check firecrawl search alexandria '<data you need>' for a ready-made workflow or data provider. Inspect a matching contract with firecrawl list <provider> <capability> --pretty and execute with firecrawl scrape --alexandria <provider>/<capability> --options '<input JSON>' if it covers the task. Use the exact provider, capability, and input fields from that contract. Continue with Agent when no suitable tool exists or the task requires autonomous navigation.
Quick start
# Extract structured data
firecrawl agent "extract all pricing tiers" --wait --json -o .firecrawl/pricing.json
# With a JSON schema for structured output
firecrawl agent "extract products" --schema '{"type":"object","properties":{"name":{"type":"string"},"price":{"type":"number"}}}' --wait --json -o .firecrawl/products.json
# Focus on specific pages
firecrawl agent "get feature list" --urls "<url>" --wait --json -o .firecrawl/features.json
Run firecrawl agent --help for the full option list.
Done when: the output file contains valid JSON answering the request — or a job ID was intentionally returned for later polling.
Job IDs
Omitting --wait returns a job ID. A UUID positional argument is auto-detected as a status check:
# Check once (equivalent to adding --status)
firecrawl agent "<job-id>"
# Wait on an existing job, polling every 10 seconds for up to 5 minutes
firecrawl agent "<job-id>" --wait --poll-interval 10 --timeout 300
# Cancel an active job
firecrawl agent "<job-id>" --cancel
Tips
- Use
--waitfor inline results; omit it only when you want a job ID to poll later (see Job IDs). - Use
--schemafor predictable, structured output — otherwise the agent returns freeform data. - Agent runs consume more credits than simple scrapes. Use
--max-creditsto cap spending. - For simple single-page extraction, prefer
scrape— it's faster and cheaper.
See also
- firecrawl-scrape — simpler single-page extraction
- firecrawl-interact — scrape + interact for manual page interaction (more control)
- firecrawl-crawl — bulk extraction without AI
- firecrawl-build-scrape — building structured extraction into an app instead of running it here
Alexandria session feedback
To report an Alexandria session outcome or a provider/capability gap, use firecrawl alexandria feedback --rating good|partial|bad --url <website> --requested-functionality '<what was needed>' --rationale '<what happened>' --json. Use observed results in the rationale. No job ID is needed; this session feedback has no job-age deadline and no credit refund. Optional --provider-feedback and --capability-feedback JSON arrays describe specific gaps; inspect firecrawl alexandria feedback --help for their fields. Use the capability issue missing_capability when a provider exists but lacks the needed capability, and new_capability_request (with requestedFunctionality) to ask for one.
Version History
-
5ec6074
Current 2026-09-27 19:59
在Alexandria反馈命令中增加了missing_capability问题类型,用于报告缺失功能。
-
dffc296
2026-09-22 08:09
新增Alexandria会话反馈命令、搜索/抓取指导、发现详情控制及合同检查路径
-
86aaf06
2026-08-27 16:50
重构文档结构,移除过时选项表与重复说明,优化触发描述以减少Token消耗,新增完成标准。
-
253abde
2026-08-19 19:06
修复了之前版本中关于CLI命令、参数及行为描述的错误声明和示例。
- 6c50c5d 2026-07-24 11:49


