Nvidia Says the Harness Beats the Model. For Browser Agents, the Harness Is Your Chrome.
Nvidia published a piece arguing that the harness, not the checkpoint, is what decides whether an agent succeeds, and it hit the Hacker News front page under the headline “Nvidia just showed that the harness, not the AI model, is now the real hero.” I read it twice. Most of the argument is about coding agents: tool definitions, retry loops, verification steps, how context gets assembled before the model ever sees a token. All reasonable. But the place this argument bites hardest isn’t a repo. It’s a browser tab.
Same weights, wildly different day
Two agents. Identical GPT-5.2, identical system prompt, identical task: check whether the Acme invoice cleared.
The first one spins up a fresh cloud Chrome. That browser has never seen your bank, your billing portal, or you. It hits an SSO wall, gets a Cloudflare interstitial for good measure, and comes back with either a polite failure or a confident guess about what was probably behind the login. The second one runs in the tab you already have open, where the invoice list is on screen and the session cookie is eleven days old and perfectly valid. Nine seconds, correct answer.
Same weights. The delta is entirely environmental, and it’s not a small delta, it’s the difference between a working feature and a demo.
What the harness actually is when the work happens on the web
For a coding agent, the harness is linters, tests, a knowledge base, structural constraints. For a browser agent it’s four things, and none of them are prompts.
Login state, first. Not credentials the agent holds, but sessions the browser already has. Second, DOM access: whether the agent reads the rendered accessibility tree of the actual page or a screenshot of an approximation of it. Third, session context, meaning the eight tabs you have open right now, which encode what you’re working on more precisely than any memory system anyone has shipped. Fourth, the identity the site sees, because a residential Chrome profile with a real history gets treated very differently from a datacenter IP running headless.
Swap the model out under those four and results shift a little. Swap those four out under a fixed model and results shift completely.
The benchmarks are grading the wrong half
Same week as the Nvidia post: a wave of open-source autoscaling browser agents, plus at least two AI browser agent leaderboards making the rounds. Every one of them benchmarks models. None of them publish what the execution environment looked like, which means a 91% and an 84% might be measuring proxy quality rather than reasoning quality, and nobody reading the table would know. We’ve written before about why those leaderboards miss real usage, and the pattern hasn’t changed since.
WebVoyager tasks run on public pages. Your work does not. The tasks that matter to you sit behind a login, and the moment login is in play, the harness dominates the score.
The autoscaling thing bugs me
Three of those launches this week bragged about spinning up a thousand parallel browsers. A thousand strangers. I have one browser and it knows who I am, which is worth more than the other 999. More on that here.
So what do you change
Not the model, mostly. If you’re on Opus 5 or GPT-5.2 or Gemini 3.1 Pro, the reasoning is not your bottleneck, and swapping between them will move your success rate by a few points while the environment moves it by forty. What actually moves is where the thing runs.
An agent in your Chrome side panel inherits every session you already have. That’s the whole design bet behind dassi: the model is whatever you point it at (your ChatGPT login, your Claude key, DeepSeek if you’re feeling thrifty), and the harness is Chrome. Not a Chrome. Yours. The one with Gmail, Salesforce, the internal admin panel that has no API and never will.
Cloud browser agents have to rebuild that from scratch on every run, and they mostly can’t, which is a problem we’ve picked at before.
The part I’m not sure about
Where this gets uncomfortable is that a great harness makes a mediocre model look sharp, and that’s fine right until you’re evaluating anything. If Dassi running Sonnet in your logged-in tab beats a frontier model in a fresh cloud browser on the same task, what did you just learn about the models? Damn near nothing. You learned about the plumbing. Nvidia’s post more or less says this out loud and then moves on, and I don’t think anyone has a clean way to separate the two yet.
Maybe that’s fine. Nobody benchmarks their office chair either, and it still determines how much work gets done.