web-browser
Controls Google Chrome via CDP to automate web navigation, form filling, and interaction.
Install
mkdir -p .claude/skills/web-browser && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/6927" && unzip -o skill.zip -d .claude/skills/web-browser && rm skill.zipInstalls to .claude/skills/web-browser
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Allows to interact with web pages by performing actions such as clicking buttons, filling out forms, and navigating links. It works by remote controlling Google Chrome or Chromium browsers using the Chrome DevTools Protocol (CDP). When Claude needs to browse the web, it can use this skill to do so.Key capabilities
- →Navigate to URLs
- →Interact with DOM elements
- →Perform device emulation
- →Execute JavaScript in browser context
- →Capture screenshots
- →Dismiss cookie dialogs
How it works
The skill uses the Chrome DevTools Protocol (CDP) to remotely control an isolated instance of Chrome or Chromium for automated interaction and data extraction.
Inputs & outputs
When to use web-browser
- →Automate website testing
- →Scrape dynamic content
- →Capture screenshots
- →Fill out complex web forms
About this skill
Web Browser Skill
Minimal CDP tools for collaborative site exploration.
Start Chrome (Prefer Headless)
./scripts/start.js --headless # Recommended: isolated reusable profile
./scripts/start.js --headless --profile # Headless with a copy of your profile
./scripts/start.js # Visible browser window when needed
./scripts/start.js --headless --reset-profile # Clear cached profile before launch
Starts Chrome with remote debugging (default port :9222). Agents should use --headless by default because it is less disruptive and supports navigation, evaluation, screenshots, emulation, and logging. User extensions are disabled for headless launches; a copied profile still provides its cookies and other browser state, but extension-based features are unavailable. Headed launches continue to load extensions normally. Use headed mode only when a person needs to see or interact with the browser, such as for pick.js, manual authentication, or debugging a headless-specific difference.
The start script only reuses a running browser when its profile and launch settings match. Close the running skill browser before switching between headless and headed mode.
Profile behavior:
- Default mode uses:
~/.cache/agent-web/browser/fresh-profile --profilemode uses:~/.cache/agent-web/browser/profile-copy- The skill does not attach to your live Chrome profile directly
- If
:9222is already used by an unknown instance, start will fail instead of reusing it
If Chrome is installed in a non-standard location, set:
BROWSER_BIN=/path/to/chrome ./scripts/start.js --headless
Optional debug endpoint override:
BROWSER_DEBUG_PORT=9333 ./scripts/start.js --headless
Navigate
./scripts/nav.js https://example.com
./scripts/nav.js https://example.com --new
Navigate current tab or open new tab.
Device Emulation (Mobile)
./scripts/emulate.js --list
./scripts/emulate.js iphone-14
./scripts/emulate.js pixel-7 --landscape
./scripts/emulate.js --reset
Set an active device emulation preference (viewport, DPR, touch, UA) for browser skill commands. Use --reset to clear.
Commands like nav.js, eval.js, pick.js, dismiss-cookies.js, and screenshot.js automatically apply the active preference.
Evaluate JavaScript
./scripts/eval.js 'document.title'
./scripts/eval.js 'document.querySelectorAll("a").length'
./scripts/eval.js 'document.querySelector("button")?.click(); "clicked"'
./scripts/eval.js 'await Promise.resolve(document.title)'
./scripts/eval.js 'JSON.stringify(Array.from(document.querySelectorAll("a")).map(a => ({ text: a.textContent.trim(), href: a.href })).filter(link => !link.href.startsWith("https://")))'
Execute JavaScript in the active tab. Input can be an expression or statement list; the console-style completion value is printed and promises/top-level await are awaited. Be careful with string escaping, best to use single quotes.
Screenshot
./scripts/screenshot.js
./scripts/screenshot.js --full-page
./scripts/screenshot.js --device iphone-14
./scripts/screenshot.js --device pixel-7 --full-page
Takes a screenshot and returns a temp file path.
- Default: current viewport
--full-page: captures full document height--device <preset>: temporary mobile emulation for that screenshot only
Pick Elements
./scripts/pick.js "Click the submit button"
Interactive element picker. Click to select, Cmd/Ctrl+Click for multi-select, Enter to finish. This requires headed Chrome; launch start.js without --headless.
Dismiss Cookie Dialogs
./scripts/dismiss-cookies.js # Accept cookies
./scripts/dismiss-cookies.js --reject # Reject cookies (where possible)
Automatically dismisses EU cookie consent dialogs.
Run after navigating to a page:
./scripts/nav.js https://example.com && ./scripts/dismiss-cookies.js
Quick Mobile Debug Flow
./scripts/start.js --headless
./scripts/nav.js https://example.com
./scripts/emulate.js iphone-14
./scripts/nav.js https://example.com # reload with mobile UA
./scripts/dismiss-cookies.js
./scripts/screenshot.js --full-page
Background Logging (Console + Errors + Network)
Automatically started by start.js and writes JSONL logs to:
~/.cache/agent-web/logs/YYYY-MM-DD/<targetId>.jsonl
Manually start:
./scripts/watch.js
Tail latest log:
./scripts/logs-tail.js # dump current log and exit
./scripts/logs-tail.js --follow # keep following
Summarize network responses:
./scripts/net-summary.js
When not to use it
- →When direct access to a user's live browser profile is required
Prerequisites
Limitations
- →Does not attach to live user profiles
- →Requires Chrome or Chromium to be installed
How it compares
It provides a scriptable, isolated environment for browser automation instead of relying on manual browser interaction.
Compared to similar skills
web-browser side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| web-browser (this skill) | 2 | 2mo | Review | Intermediate |
| browser-automation | 39 | 1mo | Review | Intermediate |
| browser-use | 64 | 3mo | Review | Intermediate |
| gui-task | 0 | 5mo | No flags | Beginner |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by mitsuhiko
View all by mitsuhiko →You might also like
browser-automation
browserbase
Automate web browser interactions using natural language via CLI commands. Use when the user asks to browse websites, navigate web pages, extract data from websites, take screenshots, fill forms, click buttons, or interact with web applications. Triggers include "browse", "navigate to", "go to website", "extract data from webpage", "screenshot", "web scraping", "fill out form", "click on", "search for on the web". When taking actions be as specific as possible.
browser-use
browser-use
Automates browser interactions for web testing, form filling, screenshots, and data extraction. Use when the user needs to navigate websites, interact with web pages, fill forms, take screenshots, or extract information from web pages.
gui-task
bivex
>
telegrab-fetch
saturov
Use when the user needs to download one Telegram video from a post link through this repository's telegrab CLI with deterministic preflight checks and machine-readable result fields.
mcporter
consuelohq
--- name: mcporter description: Call MCP servers via MCPorter - browser automation, vision, and other MCP tools homepage: https://github.com/steipete/mcporter metadata: { "openclaw": { "emoji": "🧳", "requires": { "bins": ["npx"] } } }
powerskills-browser
aloth
Edge browser automation via Chrome DevTools Protocol (CDP). List tabs, navigate, take screenshots, extract page content/HTML, execute JavaScript, click elements, type text, fill forms, scroll. Use when needing to control Edge browser, scrape web content, automate web forms, or take browser screensho