Agent Skills › nanocoai/nanoclaw › agent-browser

agent-browser

GitHub

提供浏览器自动化能力,支持网页导航、DOM快照分析、元素交互、数据提取及截图。适用于自动化测试、表单填写、信息抓取及Web应用操作场景。

container/skills/agent-browser/SKILL.md nanocoai/nanoclaw

Trigger Scenarios

需要浏览网页或访问特定URL 进行Web页面自动化测试 从网页提取结构化数据 模拟用户填写表单或点击按钮 获取网页截图或PDF

Install

npx skills add nanocoai/nanoclaw --skill agent-browser -g -y
More Options

Non-standard path

npx skills add https://github.com/nanocoai/nanoclaw/tree/main/container/skills/agent-browser -g -y

Use without installing

npx skills use nanocoai/nanoclaw@agent-browser

指定 Agent (Claude Code)

npx skills add nanocoai/nanoclaw --skill agent-browser -a claude-code -g -y

安装 repo 全部 skill

npx skills add nanocoai/nanoclaw --all -g -y

预览 repo 内 skill

npx skills add nanocoai/nanoclaw --list

SKILL.md

Frontmatter
{
    "name": "agent-browser",
    "description": "Browse the web for any task — research topics, read articles, interact with web apps, fill forms, take screenshots, extract data, and test web pages. Use whenever a browser would be useful, not just when the user explicitly asks.",
    "allowed-tools": "Bash(agent-browser:*)"
}

Browser Automation with agent-browser

Quick start

agent-browser open <url>        # Navigate to page
agent-browser snapshot -i       # Get interactive elements with refs
agent-browser click @e1         # Click element by ref
agent-browser fill @e2 "text"   # Fill input by ref
agent-browser close             # Close browser

Core workflow

  1. Navigate: agent-browser open <url>
  2. Snapshot: agent-browser snapshot -i (returns elements with refs like @e1, @e2)
  3. Interact using refs from the snapshot
  4. Re-snapshot after navigation or significant DOM changes

Commands

Navigation

agent-browser open <url>      # Navigate to URL
agent-browser back            # Go back
agent-browser forward         # Go forward
agent-browser reload          # Reload page
agent-browser close           # Close browser

Snapshot (page analysis)

agent-browser snapshot            # Full accessibility tree
agent-browser snapshot -i         # Interactive elements only (recommended)
agent-browser snapshot -c         # Compact output
agent-browser snapshot -d 3       # Limit depth to 3
agent-browser snapshot -s "#main" # Scope to CSS selector

Interactions (use @refs from snapshot)

agent-browser click @e1           # Click
agent-browser dblclick @e1        # Double-click
agent-browser fill @e2 "text"     # Clear and type
agent-browser type @e2 "text"     # Type without clearing
agent-browser press Enter         # Press key
agent-browser hover @e1           # Hover
agent-browser check @e1           # Check checkbox
agent-browser uncheck @e1         # Uncheck checkbox
agent-browser select @e1 "value"  # Select dropdown option
agent-browser scroll down 500     # Scroll page
agent-browser upload @e1 file.pdf # Upload files

Get information

agent-browser get text @e1        # Get element text
agent-browser get html @e1        # Get innerHTML
agent-browser get value @e1       # Get input value
agent-browser get attr @e1 href   # Get attribute
agent-browser get title           # Get page title
agent-browser get url             # Get current URL
agent-browser get count ".item"   # Count matching elements

Screenshots & PDF

agent-browser screenshot          # Save to temp directory
agent-browser screenshot path.png # Save to specific path
agent-browser screenshot --full   # Full page
agent-browser pdf output.pdf      # Save as PDF

Wait

agent-browser wait @e1                     # Wait for element
agent-browser wait 2000                    # Wait milliseconds
agent-browser wait --text "Success"        # Wait for text
agent-browser wait --url "**/dashboard"    # Wait for URL pattern
agent-browser wait --load networkidle      # Wait for network idle

Waiting for a custom condition — ALWAYS bound it

Prefer the built-in wait subcommands above. Only fall back to eval-polling when you must wait on a custom JS condition (e.g. a spinner disappearing or a "Send" button re-enabling in a chat UI).

Never write an unbounded wait loop. A bare until … do sleep; done that polls a page condition will loop forever if the condition never becomes true (page failed to load, selector changed, network stalled). That does not just fail the command — it wedges the entire agent turn: the runner keeps the model stream open, later messages get silently swallowed, and the container can hang for hours without the host's stuck-detection firing.

Always cap the wait with BOTH a wall-clock timeout and a max-attempts counter, and always exit the loop (never leave a sleep loop as the last thing running):

# Bounded wait: succeeds when the condition is met, gives up after ~90s.
timeout 90 bash -c '
  for i in $(seq 1 30); do
    if agent-browser eval "document.querySelector(\".loading\") === null" 2>/dev/null | grep -q true; then
      echo READY; exit 0
    fi
    sleep 3
  done
  echo TIMEOUT; exit 1
'
# Check the exit status / output: on TIMEOUT, snapshot the page and decide —
# do NOT re-enter another unbounded wait.

If the wait times out, treat it as a real failure: take a snapshot -i or screenshot to see the actual page state, report what you found, and move on. Retrying the same unbounded wait is what causes the hang.

Semantic locators (alternative to refs)

agent-browser find role button click --name "Submit"
agent-browser find text "Sign In" click
agent-browser find label "Email" fill "user@test.com"
agent-browser find placeholder "Search" type "query"

Authentication with saved state

# Login once
agent-browser open https://app.example.com/login
agent-browser snapshot -i
agent-browser fill @e1 "username"
agent-browser fill @e2 "password"
agent-browser click @e3
agent-browser wait --url "**/dashboard"
agent-browser state save auth.json

# Later: load saved state
agent-browser state load auth.json
agent-browser open https://app.example.com/dashboard

Cookies & Storage

agent-browser cookies                     # Get all cookies
agent-browser cookies set name value      # Set cookie
agent-browser cookies clear               # Clear cookies
agent-browser storage local               # Get localStorage
agent-browser storage local set k v       # Set value

JavaScript

agent-browser eval "document.title"   # Run JavaScript

Example: Form submission

agent-browser open https://example.com/form
agent-browser snapshot -i
# Output shows: textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Submit" [ref=e3]

agent-browser fill @e1 "user@example.com"
agent-browser fill @e2 "password123"
agent-browser click @e3
agent-browser wait --load networkidle
agent-browser snapshot -i  # Check result

Example: Data extraction

agent-browser open https://example.com/products
agent-browser snapshot -i
agent-browser get text @e1  # Get product title
agent-browser get attr @e2 href  # Get link URL
agent-browser screenshot products.png

Version History

  • 882305e Current 2026-08-20 09:43

Same Skill Collection

.claude/skills/add-anydoc/container-skills/convert-documents-to-markdown/SKILL.md
.claude/skills/add-anydoc/SKILL.md
.claude/skills/add-atomic-chat-tool/SKILL.md
.claude/skills/add-clidash/SKILL.md
.claude/skills/add-codex/SKILL.md
.claude/skills/add-dashboard/SKILL.md
.claude/skills/add-deltachat/SKILL.md
.claude/skills/add-dial-number/SKILL.md
.claude/skills/add-dial/SKILL.md
.claude/skills/add-emacs/SKILL.md
.claude/skills/add-imessage/SKILL.md
.claude/skills/add-iron-proxy/SKILL.md
.claude/skills/add-karpathy-llm-wiki/SKILL.md
.claude/skills/add-linear/SKILL.md
.claude/skills/add-macos-statusbar/SKILL.md
.claude/skills/add-mattermost/SKILL.md
.claude/skills/add-mnemon/SKILL.md
.claude/skills/add-ollama-provider/SKILL.md
.claude/skills/add-ollama-tool/SKILL.md
.claude/skills/add-onecli/SKILL.md
.claude/skills/add-opencode/SKILL.md
.claude/skills/add-resend/SKILL.md
.claude/skills/add-rtk/SKILL.md
.claude/skills/add-signal/SKILL.md
.claude/skills/add-tavily-tool/SKILL.md
.claude/skills/add-teams/SKILL.md
.claude/skills/add-vercel/container-skills/vercel-cli/SKILL.md
.claude/skills/add-vercel/SKILL.md
.claude/skills/add-wechat/SKILL.md
.claude/skills/add-whatsapp-cloud/SKILL.md
.claude/skills/add-whatsapp/SKILL.md
.claude/skills/customize/SKILL.md
.claude/skills/debug/SKILL.md
.claude/skills/init-first-agent/SKILL.md
.claude/skills/init-onecli/SKILL.md
.claude/skills/manage-channels/SKILL.md
.claude/skills/manage-mounts/SKILL.md
.claude/skills/migrate-from-openclaw/SKILL.md
.claude/skills/migrate-from-v1/SKILL.md
.claude/skills/migrate-memory/SKILL.md
.claude/skills/migrate-nanoclaw/SKILL.md
.claude/skills/migrate-slack-agents/SKILL.md
.claude/skills/setup/SKILL.md
.claude/skills/slack-a2a-rooms/SKILL.md
.claude/skills/slack-agent-flow/SKILL.md
.claude/skills/update-nanoclaw/SKILL.md
.claude/skills/update-skills/SKILL.md
container/skills/frontend-engineer/SKILL.md
container/skills/onecli-gateway/SKILL.md

Metadata

Files
0
Version
5e4f4b7
Hash
8b7674ba
Indexed
2026-08-20 09:43

ホーム - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-29 02:00
浙ICP备14020137号-1