Open source · Rust · Apache-2.0 or MIT
A headless-first browser for AI agents.
CatPaw runs pages' JavaScript and DOM the way a browser does, and gives the agent what it needs: a compact snapshot with stable references, what changed after each action, and a clear signal when the page has settled. It is written from scratch in Rust, not wrapped around Chromium.
curl -fsSL https://catpaw.sh/install.sh | shirm https://catpaw.sh/install.ps1 | iexVersion 0.1 is a preview. The installer shows what it will write and asks first. What works and what does not yet.
Why
Built for the agent, not adapted to it
CatPaw is not a Chromium wrapper, and not a rendering engine with an automation API added on. Its first user is an agent. It runs JavaScript and the DOM faithfully, computes layout only when something asks for geometry, and answers the questions an agent has: what is on the page, what changed, and whether the page is done.
-
Compact snapshots, stable refs
snapshotreturns a CatPaw Snapshot Text tree, a superset of Playwright's aria snapshot. Elements are named by refs such ase12that stay valid until the element leaves the page and are never reused. A whole snapshot is capped at 4000 tokens; the rest is folded for the agent to open. -
Diffs after actions
An action answers once the page has settled, with what happened and what changed, so one agent step is one round trip. A new document comes back whole.
-
Settledness
CatPaw knows every pending fetch, timer, animation frame and microtask. A page that does not settle is reported with its causes, down to the script position that started them. Time a page spends only on timers passes at once.
-
Confirmations before side effects
Under the default policy, a form submission or an upload waits for the user. A host that supports MCP elicitation asks there; otherwise CatPaw opens an approval page in the user's own browser. The agent never sees the key.
-
Hand-off to the user
When a site needs a person, such as a login or a check meant for humans,
handoffopens the tab in the user's own browser. They click, type and scroll as on the page itself, and the agent gets the page back with what they typed masked. -
Flight recorder
--flight-logkeeps a journal of every call, confirmation and decision, with a screenshot per action if you ask for one. Typed passwords are kept as their length. -
Record and replay
--record-harkeeps a session's traffic and--replay-harserves it back with no network. With a fixed seed and clock, a replay gives the same results byte for byte. -
Honest identity
CatPaw says what it is, signs requests with Web Bot Auth when you give it a key, and ships no fingerprint impersonation and no CAPTCHA solvers. Loopback and private addresses are refused unless you allow them. For site owners.
ok click e16 button "Add to cart"
# s4 diff-from=s3 tab=t1 doc=d1 url=(same) scroll=0,0 settled=yes changed=1 added=1 removed=1 unchanged=27
~ e11 button "Cart, empty" → "Cart, 1 items"
+ e37 button "Remove" (in e13, after e15)
- e16 button "Add to cart"
Measured
What an agent reads
Twenty tasks on sites made for automation practice live in the repository, each with its recorded traffic and the transcript an agent sees. The table sets the bytes of every tool result an agent receives against Playwright MCP taking the same steps.
| Task | CatPaw | Playwright MCP |
|---|---|---|
| Practice-site tasks, all but one (see below) | 71 calls 68099 (~19457) | 118 calls 243282 (~69509) |
| Hacker News, second page | 2 calls 14618 (~4177) | 4 calls 98113 (~28032) |
| Wikipedia, search to an article | 2 calls 11224 (~3207) | 4 calls 569358 (~162674) |
| Tool list, sent on every turn | 12.1 KB | 20.3 KB |
How this was measured: bytes are all the tool results an agent receives over a task. CatPaw's numbers come from the recordings. Playwright MCP is @playwright/mcp 0.0.83 with headless Chrome, taking the same steps live on 2026-10-08, the median of three runs. Playwright MCP keeps the page snapshot in a file and links it from a result when the page changed; an agent reads it to see the page and find its next target, so the file counts too, as one more call. CatPaw's numbers include the calls that wait for the user's approval, which Playwright MCP does not make. The task in which the user declines has no counterpart there, so the totals leave it out. The content-site rows were measured the same way on the same day; those recordings stay out of the repository. The per-task table and the commands that regenerate it are in the README.
Install
One file, and you are asked first
curl -fsSL https://catpaw.sh/install.sh | shirm https://catpaw.sh/install.ps1 | iexThe installer downloads the release for your machine from GitHub Releases and checks it against the release's SHA256SUMS. It then shows the file it will write, and whether it replaces one, and waits for your yes. Nothing is written before that. Running it again upgrades in place. You can read install.sh and install.ps1 first.
What it writes
- macOS and Linux: one file,
~/.catpaw/bin/catpaw. No sudo, and your shell profile and PATH are left alone; it prints the line to add if the directory is not on your PATH. - Windows: one file,
%LOCALAPPDATA%\Programs\CatPaw\catpaw.exe, and that directory is added to your user PATH. No administrator rights.
Releases are built for Linux on x86_64 and aarch64 (glibc), macOS on Apple silicon, and Windows on x86_64; Windows 11 on ARM runs the x64 build under emulation. On an Intel Mac, build from source.
Options
Build from source
git clone https://github.com/KernelErr/CatPaw
cd CatPaw
cargo build --release -p catpawThis needs Rust 1.89 or later and a C compiler. The binary is target/release/catpaw.
Connect
Give it to your agent
catpaw mcp --stdio serves the browser to an agent over the Model Context Protocol. catpaw setup registers it with your agent host:
catpaw setup claude-code # prints the claude mcp add command
catpaw setup claude-code --write # or writes .mcp.json in this project
catpaw setup codex --write # adds it to ~/.codex/config.toml
catpaw setup cursor --write # adds it to ~/.cursor/mcp.jsonWithout --write, setup prints what to run or add. Options after -- go to catpaw mcp, for example catpaw setup cursor -- --policy strict, which also asks before scripts send data to other sites and before evaluate.
The tools are navigate, snapshot, click, type, fill, press, select, act, wait, read, screenshot, evaluate, tabs, logs and handoff. Windows a page opens become tabs. Errors say what to try next.
To look at a page yourself, without an agent:
catpaw fetch https://example.com --js --snapshotStatus
A preview, plainly
Version 0.1 is a preview. The first four milestones are done: fetch and read, scripts, interaction, and the agent API. Expect gaps, and expect some sites not to work.
What works
- JavaScript on the Boa engine: classic and module scripts, the DOM, events,
fetchand XHR, storage, workers, WebSocket, Canvas 2D and Web Crypto. React, Vue, Svelte, Lit, htmx and Alpine sites run. - Layout on demand, and screenshots of backgrounds, borders, text and form controls.
- Trusted pointer and keyboard input, forms, navigation and history, iframes, and popups as tabs.
- The MCP server with confirmations, hand-off, the flight journal, profiles and HAR record and replay. Twenty practice tasks replay in CI, twice, byte for byte.
Not yet
- There is no JIT. Boa interprets, so a single-page app's own scripts run many times slower than in Chrome.
- Screenshots do not draw images, gradients or rounded corners.
- No media or WebAssembly. Tables are not laid out as a grid, and images have no intrinsic size.
- Style sheets do not see runtime state:
:hover,:focusand:checkedafter a click are not matched. - Downloads stay in memory, not on disk. Recordings redact secrets in requests, cookies, credential headers and JSON answers, not in HTML or other text.
- CatPaw will not pass every site. Whether to let it in is the site owner's decision.
Next on the roadmap: challenge detection and measured pass rates, then multi-tenant limits, OpenTelemetry, Docker, a CDP subset for Puppeteer and an optional V8 backend. The full list of known gaps is in the architecture notes.
Read more
Documentation
- README: features, the agent tools, confirmations and hand-off, the measurements.
- Architecture: how the engine is put together, and its known gaps.
- Decision records, among them identity and challenges and the agent protocol.
- For site owners: what CatPaw traffic is, and how to block or allow it.
- Questions and bugs: GitHub issues. Security reports: a private advisory.