Open source · Rust · Apache-2.0 or MIT

A headless-first browser for AI agents.

CatPaw runs pages' JavaScript and DOM the way a browser does, and gives the agent what it needs: a compact snapshot with stable references, what changed after each action, and a clear signal when the page has settled. It is written from scratch in Rust, not wrapped around Chromium.

macOS and Linux
curl -fsSL https://catpaw.sh/install.sh | sh
Windows (PowerShell)
irm https://catpaw.sh/install.ps1 | iex

Version 0.1 is a preview. The installer shows what it will write and asks first. What works and what does not yet.

Why

Built for the agent, not adapted to it

CatPaw is not a Chromium wrapper, and not a rendering engine with an automation API added on. Its first user is an agent. It runs JavaScript and the DOM faithfully, computes layout only when something asks for geometry, and answers the questions an agent has: what is on the page, what changed, and whether the page is done.

  • Compact snapshots, stable refs

    snapshot returns a CatPaw Snapshot Text tree, a superset of Playwright's aria snapshot. Elements are named by refs such as e12 that stay valid until the element leaves the page and are never reused. A whole snapshot is capped at 4000 tokens; the rest is folded for the agent to open.

  • Diffs after actions

    An action answers once the page has settled, with what happened and what changed, so one agent step is one round trip. A new document comes back whole.

  • Settledness

    CatPaw knows every pending fetch, timer, animation frame and microtask. A page that does not settle is reported with its causes, down to the script position that started them. Time a page spends only on timers passes at once.

  • Confirmations before side effects

    Under the default policy, a form submission or an upload waits for the user. A host that supports MCP elicitation asks there; otherwise CatPaw opens an approval page in the user's own browser. The agent never sees the key.

  • Hand-off to the user

    When a site needs a person, such as a login or a check meant for humans, handoff opens the tab in the user's own browser. They click, type and scroll as on the page itself, and the agent gets the page back with what they typed masked.

  • Flight recorder

    --flight-log keeps a journal of every call, confirmation and decision, with a screenshot per action if you ask for one. Typed passwords are kept as their length.

  • Record and replay

    --record-har keeps a session's traffic and --replay-har serves it back with no network. With a fixed seed and clock, a replay gives the same results byte for byte.

  • Honest identity

    CatPaw says what it is, signs requests with Web Bot Auth when you give it a key, and ships no fingerprint impersonation and no CAPTCHA solvers. Loopback and private addresses are refused unless you allow them. For site owners.

What an agent reads after clicking "Add to cart": the action, a header, and the three lines that changed.
ok click e16 button "Add to cart"
# s4 diff-from=s3 tab=t1 doc=d1 url=(same) scroll=0,0 settled=yes changed=1 added=1 removed=1 unchanged=27
~ e11 button "Cart, empty" → "Cart, 1 items"
+ e37 button "Remove" (in e13, after e15)
- e16 button "Add to cart"

Measured

What an agent reads

Twenty tasks on sites made for automation practice live in the repository, each with its recorded traffic and the transcript an agent sees. The table sets the bytes of every tool result an agent receives against Playwright MCP taking the same steps.

Bytes, with tokens estimated at 3.5 bytes each
TaskCatPawPlaywright MCP
Practice-site tasks, all but one (see below)71 calls
68099 (~19457)
118 calls
243282 (~69509)
Hacker News, second page2 calls
14618 (~4177)
4 calls
98113 (~28032)
Wikipedia, search to an article2 calls
11224 (~3207)
4 calls
569358 (~162674)
Tool list, sent on every turn12.1 KB20.3 KB

How this was measured: bytes are all the tool results an agent receives over a task. CatPaw's numbers come from the recordings. Playwright MCP is @playwright/mcp 0.0.83 with headless Chrome, taking the same steps live on 2026-10-08, the median of three runs. Playwright MCP keeps the page snapshot in a file and links it from a result when the page changed; an agent reads it to see the page and find its next target, so the file counts too, as one more call. CatPaw's numbers include the calls that wait for the user's approval, which Playwright MCP does not make. The task in which the user declines has no counterpart there, so the totals leave it out. The content-site rows were measured the same way on the same day; those recordings stay out of the repository. The per-task table and the commands that regenerate it are in the README.

Install

One file, and you are asked first

macOS and Linux
curl -fsSL https://catpaw.sh/install.sh | sh
Windows (PowerShell)
irm https://catpaw.sh/install.ps1 | iex

The installer downloads the release for your machine from GitHub Releases and checks it against the release's SHA256SUMS. It then shows the file it will write, and whether it replaces one, and waits for your yes. Nothing is written before that. Running it again upgrades in place. You can read install.sh and install.ps1 first.

What it writes

  • macOS and Linux: one file, ~/.catpaw/bin/catpaw. No sudo, and your shell profile and PATH are left alone; it prints the line to add if the directory is not on your PATH.
  • Windows: one file, %LOCALAPPDATA%\Programs\CatPaw\catpaw.exe, and that directory is added to your user PATH. No administrator rights.

Releases are built for Linux on x86_64 and aarch64 (glibc), macOS on Apple silicon, and Windows on x86_64; Windows 11 on ARM runs the x64 build under emulation. On an Intel Mac, build from source.

Options

Another directory
curl -fsSL https://catpaw.sh/install.sh | sh -s -- --dir ~/bin
& ([scriptblock]::Create((irm https://catpaw.sh/install.ps1))) -Dir 'D:\Tools\CatPaw'
Without questions, for scripts and CI

Add --yes (sh) or -Yes (PowerShell), or set CATPAW_YES=1. Without a terminal to ask on and without one of these, the installer stops and says so.

A particular release

--version v0.1.0 or -Version v0.1.0, or set CATPAW_VERSION. CATPAW_INSTALL_DIR does the same as --dir; a flag wins over its variable.

Uninstall
curl -fsSL https://catpaw.sh/install.sh | sh -s -- --uninstall
& ([scriptblock]::Create((irm https://catpaw.sh/install.ps1))) -Uninstall

It removes the binary (and on Windows the PATH entry), asks separately before deleting CatPaw's approval key, and lists what to remove from your agent hosts. Pass the same --dir or -Dir if you used one. By hand: delete the binary, and on Windows remove its directory from your user PATH.

Build from source

git clone https://github.com/KernelErr/CatPaw
cd CatPaw
cargo build --release -p catpaw

This needs Rust 1.89 or later and a C compiler. The binary is target/release/catpaw.

Connect

Give it to your agent

catpaw mcp --stdio serves the browser to an agent over the Model Context Protocol. catpaw setup registers it with your agent host:

catpaw setup claude-code          # prints the claude mcp add command
catpaw setup claude-code --write  # or writes .mcp.json in this project
catpaw setup codex --write        # adds it to ~/.codex/config.toml
catpaw setup cursor --write       # adds it to ~/.cursor/mcp.json

Without --write, setup prints what to run or add. Options after -- go to catpaw mcp, for example catpaw setup cursor -- --policy strict, which also asks before scripts send data to other sites and before evaluate.

The tools are navigate, snapshot, click, type, fill, press, select, act, wait, read, screenshot, evaluate, tabs, logs and handoff. Windows a page opens become tabs. Errors say what to try next.

To look at a page yourself, without an agent:

catpaw fetch https://example.com --js --snapshot

Status

A preview, plainly

Version 0.1 is a preview. The first four milestones are done: fetch and read, scripts, interaction, and the agent API. Expect gaps, and expect some sites not to work.

What works

  • JavaScript on the Boa engine: classic and module scripts, the DOM, events, fetch and XHR, storage, workers, WebSocket, Canvas 2D and Web Crypto. React, Vue, Svelte, Lit, htmx and Alpine sites run.
  • Layout on demand, and screenshots of backgrounds, borders, text and form controls.
  • Trusted pointer and keyboard input, forms, navigation and history, iframes, and popups as tabs.
  • The MCP server with confirmations, hand-off, the flight journal, profiles and HAR record and replay. Twenty practice tasks replay in CI, twice, byte for byte.

Not yet

  • There is no JIT. Boa interprets, so a single-page app's own scripts run many times slower than in Chrome.
  • Screenshots do not draw images, gradients or rounded corners.
  • No media or WebAssembly. Tables are not laid out as a grid, and images have no intrinsic size.
  • Style sheets do not see runtime state: :hover, :focus and :checked after a click are not matched.
  • Downloads stay in memory, not on disk. Recordings redact secrets in requests, cookies, credential headers and JSON answers, not in HTML or other text.
  • CatPaw will not pass every site. Whether to let it in is the site owner's decision.

Next on the roadmap: challenge detection and measured pass rates, then multi-tenant limits, OpenTelemetry, Docker, a CDP subset for Puppeteer and an optional V8 backend. The full list of known gaps is in the architecture notes.

Read more

Documentation