browser

GitHub

通过 omowright 驱动真实浏览器,支持用户已登录环境(attached)和代码启动环境(owned),用于表单交互、截图、QA 测试及自动化操作。

packages/shared-skills/skills/browser/SKILL.md code-yeongyu/oh-my-openagent

Trigger Scenarios

需要与已登录网站交互 执行网页自动化测试 生成网页截图 处理表单填写与点击

Install

npx skills add code-yeongyu/oh-my-openagent --skill browser -g -y
More Options

Non-standard path

npx skills add https://github.com/code-yeongyu/oh-my-openagent/tree/dev/packages/shared-skills/skills/browser -g -y

Use without installing

npx skills use code-yeongyu/oh-my-openagent@browser

指定 Agent (Claude Code)

npx skills add code-yeongyu/oh-my-openagent --skill browser -a claude-code -g -y

安装 repo 全部 skill

npx skills add code-yeongyu/oh-my-openagent --all -g -y

预览 repo 内 skill

npx skills add code-yeongyu/oh-my-openagent --list

SKILL.md

Frontmatter
{
    "name": "browser",
    "description": "Drives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch."
}

Browser

One library, two engines. omowright ships inside this skill; choose the engine before you act:

You need Engine Entry point
A site the user is signed into, their open tabs, a form, a click-through, a screenshot, web QA, an extension popup attached — the user's own browser through BrowserSkill connectBrowserSkill()
A throwaway profile, bot-scoring evasion, a CAPTCHA, network interception, a QA flight trace, coordinate control, headless runs owned — a browser your code launches connectPipe() / connectCloakProfile() — references/owned-engine/README.md
Text out of a URL, a 403 bypass, a platform that blocks fetchers neither the ultimate-browsing skill

Attached is the default, because it is the only engine carrying the user's logins and the only one where a human is a single call away. Never substitute one engine for the other silently: if the attached engine is not set up, run the onboarding script and tell the user its one remaining step.

Step 0 — load omowright and prove the stack

const { loadOmowright } = await import("<skill-root>/scripts/omowright.mjs")
const { omowright } = await loadOmowright()          // { connectBrowserSkill, bskSnapshot, connectPipe, ... }
node "<skill-root>/scripts/browser-doctor.mjs" --json
State Meaning Next
ready CLI, daemon and a connected browser start a session
no-cli / no-daemon / no-extension something is missing node "<skill-root>/scripts/browser-install.mjs" [--browser=<id>] prepares everything it can for the browser the user uses, then prints the single step only the user can do (relaunch that browser and click Enable); relay it verbatim, wait, re-run the doctor
choose-browser the signals do not single out one browser (Safari/Firefox default, an idle default while another browser runs, several in use) nothing was installed; take the browser from memory or ask the user, then browser-install.mjs --browser=<id>
no-browser-support no Chromium-family profile on this machine say so and stop

Install into the browser the user actually uses, never into whatever happens to be on disk. Before installing, check your memory for the user's browser; otherwise read the doctor's browser (picked from the OS default browser, running apps and recent use — candidates shows the evidence). If memory and the doctor disagree, or the doctor says choose-browser, ask the user. Pass the answer as --browser=<id> and record it in memory. A Chrome that is merely installed is not their browser.

Never launch a headless browser because the attached one is missing. It has none of the user's sessions, so every login turns into a ladder you should not be climbing. Say which state you hit and ask.

The loop (attached)

const session = await omowright.connectBrowserSkill({ name: "<task>", focused: false })
try {
  await session.navigate("https://example.com/", { waitUntil: "load" })
  const { tree, refs, css } = await omowright.bskSnapshot(session, { interactive: true })  // OmOWright tree + refs, no trace in the page
  await session.click({ selector: css.e3 })                                                 // css[ref] is null inside shadow roots:
  const vom = await session.observe({ maxTokens: 4000 })                                    //   then read the daemon's own tree ...
  await session.click("@e7")                                                                //   ... and click its @eN ref
  await session.fill(css.e5, "hello")
  await session.press("Enter")
  await session.waitForNavigation({ waitUntil: "load" })
  const shot = await session.screenshot()                                                   // { buffer, width, height, captureId }
} finally {
  await session.stop()                                                                      // success AND failure; returns borrowed tabs
}
  1. Read before every action. bskSnapshot refs and observe @eN refs are reissued on each call; use a ref in the same cycle you read it.
  2. Navigation and large DOM changes stale every ref. Read again rather than reusing.
  3. Two identical failures mean change approach, not retry. A third identical attempt is a defect.
  4. Borrow a user tab explicitly (tabList({ scope: "user" }), tabBorrow(id), tabReturn(id)). Borrowing prompts the user; never invent tab ids and never repeat a denied borrow.
  5. Always stop() the session, on success and on failure.

Every method, its options, and the failure codes are in references/commands.md.

When a human is the only way through

Login, CAPTCHA, OTP, a payment confirmation, a consent dialog:

const outcome = await session.requestHelp({ prompt: "<what you need done>", targets: ["@e4"], timeoutMs: 300_000 })

Then read the page again. Respect a cancelled or timed_out outcome; do not work around it by changing the extension's automation settings.

Rules

  • Never read credentials through the page. No evaluate that extracts a password, token, cookie or recovery code. The value of the attached engine is that the browser is already signed in.
  • Never clear cookies, cache or site data. It is the user's real profile; clearing it logs them out everywhere. No flow here needs it.
  • focused: false by default. The browser belongs to someone who is probably using it.
  • One short, named session per task, always stopped.
  • Bot-scored or WAF targets go to the owned engine. The attached engine's daemon enables console capture on every tab it drives, which is a known automation signal; CloakBrowser through connectCloakProfile() is the stealth path.

Where the rest lives

Topic Read
Session methods, targets, options, error codes references/commands.md
Installing: CLI, daemon, extension, the one human step, blocklisted extension references/install.md
Agent on one machine, browser on another references/remote.md
Owned engine: launch, snapshot ladder, network, frames, human handoff references/owned-engine/README.md
Reading a 1Password vault the user has unlocked references/recipes/1password.md

Version History

  • d83d692 Current 2026-09-28 12:20

    修复仅向用户实际使用的浏览器安装 BrowserSkill,新增 choose-browser 状态以解决多浏览器识别问题。

  • d862a09 2026-09-23 01:20

Same Skill Collection

.agents/skills/get-unpublished-changes/SKILL.md
.agents/skills/github-triage/SKILL.md
.agents/skills/omomomo/SKILL.md
.agents/skills/publish/SKILL.md
.agents/skills/remove-deadcode/SKILL.md
.opencode/skills/github-triage/SKILL.md
packages/omo-codex/plugin/skills/init-deep/SKILL.md
packages/omo-senpi/plugin/skills/init-deep/SKILL.md
packages/omo-senpi/skills/dag-library/SKILL.md
packages/omo-senpi/skills/give-me-tips/SKILL.md
packages/omo-senpi/skills/hyperplan/SKILL.md
packages/omo-senpi/skills/init-deep/SKILL.md
packages/omo-senpi/skills/mass-ulw/SKILL.md
packages/omo-senpi/skills/onboarding/SKILL.md
packages/omo-senpi/skills/ultrawork/SKILL.md
packages/omo-senpi/skills/ulw-loop/SKILL.md
packages/omo-senpi/skills/ulw-plan/SKILL.md
packages/omo-senpi/skills/ulw-research/SKILL.md
packages/pi-goal/SKILL.md
packages/shared-skills/skills/ast-grep/SKILL.md
packages/shared-skills/skills/coding-agent-sessions/SKILL.md
packages/shared-skills/skills/data-scientist/SKILL.md
packages/shared-skills/skills/debugging/SKILL.md
packages/shared-skills/skills/frontend/SKILL.md
packages/shared-skills/skills/git-master/SKILL.md
packages/shared-skills/skills/init-deep/SKILL.md
packages/shared-skills/skills/lsp-setup/SKILL.md
packages/shared-skills/skills/programming/SKILL.md
packages/shared-skills/skills/refactor/SKILL.md
packages/shared-skills/skills/remove-ai-slops/SKILL.md
packages/shared-skills/skills/review-work/SKILL.md
packages/shared-skills/skills/start-work/SKILL.md
packages/shared-skills/skills/ultimate-browsing/SKILL.md
packages/shared-skills/skills/ulw-execute/SKILL.md
packages/shared-skills/skills/ulw-plan/SKILL.md
packages/shared-skills/skills/ulw-research/SKILL.md
packages/shared-skills/skills/visual-qa/SKILL.md
.agents/skills/codex-qa/SKILL.md
.agents/skills/hyperplan/SKILL.md
.agents/skills/opencode-qa/SKILL.md
.agents/skills/pre-publish-review/SKILL.md
.agents/skills/security-research/SKILL.md
.agents/skills/senpi-qa/SKILL.md
.agents/skills/tech-debt-audit/SKILL.md
.agents/skills/work-with-pr/SKILL.md
.opencode/skills/hyperplan/SKILL.md
.opencode/skills/pre-publish-review/SKILL.md
.opencode/skills/work-with-pr/SKILL.md
packages/omo-codex/plugin/skills/ulw-plan/SKILL.md

Metadata

Files
0
Version
d83d692
Hash
e93740ad
Indexed
2026-09-23 01:20

ホーム - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-28 17:41
浙ICP备14020137号-1