DOM, Screenshot, or Accessibility Tree: How Browser Agents Actually Read a Page
Yamak open-sourced its browser agent this morning, an autoscaling agent landed on Show HN a few hours after that, and the AI Browser Agent Leaderboard managed to appear twice in the top five. All of it is about accuracy numbers. None of it is about the input.
By input I mean the thing the model literally reads before it decides where to click. Screenshot, raw DOM, or accessibility tree. That single choice does more to determine whether a run succeeds than whichever frontier model you’ve pointed at the task, and almost nobody writes it down in their README.
Screenshots look right until the page moves
Pixels are seductive because they match how you’d explain the task to a person. Send an image, ask for coordinates, click at (612, 344). Demos love this.
But coordinates are a claim about a page that stopped existing the instant the screenshot was taken. A lazy-loaded banner pushes the form down eighty pixels. A CSS transition is halfway through. A fetch resolves and React swaps the whole list. The model’s answer was correct for a frame that’s already gone, and the failure mode isn’t an exception you can catch, it’s a confident click on the button next to the one you wanted.
Scroll compounds it. Everything below the fold needs its own screenshot-reason-scroll cycle, each one a fresh vision call, and by the fourth cycle you’ve spent real money to look at a page a human would have skimmed in two seconds.
The DOM is technically complete and practically a firehose
Serialize a live Gmail tab and you get something in the hundreds of kilobytes: nested divs six levels deep, inline style attributes, framework bookkeeping props, tracking wrappers, SVG path data for every icon in the toolbar. All of it is faithfully what the browser holds in memory, and roughly two percent of it has any bearing on the question of which element is the reply button.
You can strip it. Everyone does. You drop <script> and <style>, prune attributes down to an allowlist, collapse text nodes. And what you have after that pruning pass is no longer the DOM, it’s a lossy summary you invented, with your own guesses baked in about which attributes mattered. Half the agent frameworks on GitHub are, underneath the prompt engineering, someone’s particular opinion about which HTML attributes to throw away.
So the accessibility tree wins by default
Chrome already computes one. It’s the structure screen readers consume, and it’s derived from the rendered page rather than the source: roles, names, states, values. A button is button "Send". A checked box is checkbox "Remember me" checked. Hidden elements are gone, decorative junk is gone, and what’s left is a few kilobytes of semantics instead of a few hundred kilobytes of markup.
It’s also stable across re-renders in a way coordinates never are, because “the button named Send” survives a layout shift that would have invalidated (612, 344).
Where the tree comes from decides everything
This is the part the leaderboards flatten out. An accessibility tree computed inside a headless Chromium on someone’s server is a tree of a logged-out page. Your Salesforce dashboard, from that machine, is a login form. Your Gmail is a marketing page about Gmail. The agent perceives that shell perfectly and reasons about it correctly, and it’s still describing a room you’ve never been in.
An extension reading the tree from your actual tab gets the post-JavaScript, post-authentication render: your inbox, your rows, your org’s custom fields. Same algorithm, completely different content. Dassi sits in the Chrome side panel for exactly this reason, and it’s why cloud browser agents can’t see your tabs no matter how good their perception layer is.
Where it still falls over
Canvas apps. Figma, most charting libraries, anything drawing to a bitmap: the tree says canvas and stops. Screenshots are genuinely the only option there.
Then there’s div soup. <div onclick> styled to look like a button has no role, no accessible name, nothing. The tree reports a generic container and the agent has to infer from surrounding text.
Shadow DOM is the annoying one, because closed shadow roots hide their internals from everything, and a page built out of web components can present as a nearly empty tree while looking completely normal to you.
Which is why the honest answer is a hybrid: tree first, screenshot when the tree comes back thin. Most serious agents do this. Almost none of them say so, and the leaderboards score the composite without ever separating out how much of a failure was the model versus how much was a page that refused to describe itself. Accuracy scores hide a lot.
If you want to watch the perception layer work on pages you’re already logged into, Dassi is in the Chrome Web Store and free. Bring your own key or sign in with ChatGPT.
Anyway. Half a decade of accessibility advocacy got ignored by most frontend teams, and now the machines are the ones filing the bug reports. Damn funny way for that debt to come due.