I counted the docker mentions in Yamak’s README this morning. Seven. Skyvern 2.0’s quickstart has nine, plus a docker-compose.yml that pulls headless Chrome, a Postgres database, and a proxy service. Browser Use needs a Python environment with Playwright. And the new AI Browser Agent Leaderboard that showed up on Hacker News alongside both of them? Every single entrant in the top ten runs on server infrastructure.

This pattern keeps repeating and I am starting to think nobody in the open-source browser agent community has stopped to ask whether it actually makes sense.

The infrastructure tax

Skyvern 2.0 scored 85.8% on WebVoyager, which is genuinely impressive, and I wrote about what those benchmark numbers mean in a previous post. But to run it yourself you need Docker, a Postgres instance, environment variables for your LLM keys, a proxy configuration, and enough patience to debug port conflicts when something inevitably collides with whatever else you have running on localhost. Yamak, the Kotlin-based newcomer with single-digit GitHub stars, ships as a desktop app but still expects you to configure a headless browser runtime and wire up API credentials through config files that look like they were designed by someone who thinks YAML is a love language.

Browser Use is maybe the most approachable of the bunch, and “approachable” in this context means “you only need Python 3.11+ and Playwright and an afternoon.”

So the people who would benefit most from browser automation — the ones drowning in repetitive tab-switching, form-filling, email-drafting work — are exactly the people least likely to survive the setup process. I keep making this point because it keeps being true.

Headless Chrome does not know you

And the server problem is not just about installation friction. It is about what headless Chrome fundamentally cannot do.

A headless browser instance running in a Docker container has no cookies, no saved passwords, no logged-in sessions, nothing. It’s a blank slate every single time, which means the first thing any server-based agent has to do is authenticate, and authentication on the modern web is a damn obstacle course of OAuth flows, MFA prompts, CAPTCHA challenges, and session tokens that expire in ways nobody fully documents. The agent spends half its energy just getting through the front door of whatever site you pointed it at, before it can even attempt the task you actually care about.

Your real browser already solved this problem. You logged into Gmail this morning. You are still logged into your company’s HR portal from last week. Your CRM has a persistent session cookie that will outlive most of your professional relationships. All of that context, all of those authenticated sessions, sitting right there in the browser you already have open.

A Chrome extension just runs

Dassi lives in Chrome’s side panel. It sees the page you see, uses the sessions you already have, and talks directly to your LLM provider with your own API key. No Docker. No Postgres. No headless anything. You install it from the Chrome Web Store and it works, which should not feel revolutionary but apparently is.

The architectural difference matters beyond convenience. When an agent runs in your actual browser, it inherits your entire authenticated state without ever needing to store or transmit your credentials. Server-based agents need your login information sent somewhere, somehow, and that somewhere is a threat surface that keeps proving the skeptics right.

So why does everyone keep building servers?

Benchmarks. That is the short answer, and it explains almost everything about the current landscape.

WebVoyager tasks run in controlled environments where headless Chrome is perfectly adequate because the benchmark does not test against real user sessions. Nobody on the leaderboard is measuring “can this agent draft a reply in the Gmail account you are actually logged into right now,” because that task is impossible to standardize, and the things that are impossible to standardize are precisely the things people do all day at work. Optimizing for leaderboard performance and optimizing for real-world usefulness have started diverging, and the gap is getting wider every month.

Open-source browser agents are doing important work. Skyvern’s Planner-Actor-Validator architecture is legitimately clever engineering. But the assumption baked into every one of these projects — that you need a server between the user and the web — is worth questioning more than it gets questioned.

The browser is already running. The sessions are already there. Maybe the server was never the right abstraction.