Someone on Hacker News last week complained that browser agent demos always use the same three canned websites, and I thought yeah, that is the whole damn problem with how people evaluate these things. So we ran dassi against the full WebVoyager benchmark from He et al. at Zhejiang University. 643 tasks across 15 live websites, real Chrome, no mocks, no sandboxes.

93.2% answer accuracy on tasks where it returned a definitive answer.

The benchmark, briefly

WebVoyager sends an agent to real websites and asks it to do things a person would do. Shopping on Amazon, pulling paper abstracts from ArXiv, checking BBC News headlines, looking up definitions on Cambridge Dictionary, querying Wolfram Alpha. The tasks range from trivial lookups that take twenty seconds to multi-step research jobs where the agent has to navigate filters, sort results, compare listings, and synthesize information from pages that Amazon keeps rearranging on a weekly basis because apparently no layout is ever good enough for them.

I spent most of last week watching the runs in a terminal and it was weirdly hypnotic. Like watching someone else play a video game, except the player is an LLM and the game is “figure out where Allrecipes hid the recipe URL this time.”

93.2%

Most failures were not actually wrong answers. The agent would return a truncated product title, or give three recipe names when the task expected a link. Formatting mismatches between what the agent found and what the evaluation regex wanted. So the information was there, sitting in the agent’s output, but wrapped in enough extra context that the grader could not parse it cleanly.

We tightened the extraction patterns after the first batch and that helped. But there is still a gap between the agent knowing an answer and expressing it in the exact format a mechanical grader accepts, which is a problem I suspect every team running these benchmarks quietly deals with.

Some tasks hit the three-minute timeout. ArXiv was the worst offender because academic paper pages are dense, navigation is deeply nested, and the agent sometimes went down rabbit holes parsing LaTeX abstracts when it should have just grabbed the title and moved on.

Why live websites make this hard

Every AI company publishes benchmark numbers, and most of those numbers come from sanitized static datasets where the right answer was locked in months ago. Browser benchmarks are a different animal. The Amazon bestseller list from last Tuesday is not today’s. BBC News rotates headlines hourly. ArXiv gets hundreds of new submissions every single day.

And then there are cookie consent banners. I did not expect cookie banners to be a meaningful source of failures but they were, which tells you something about the state of the modern web that I find both funny and a little bleak. Layout changes, CAPTCHAs from hitting a site too many times on one IP, sponsored results injecting themselves into product listings. None of this was mocked out. The agent had to cope with whatever the website threw at it on that particular Tuesday afternoon.

Speed, and what is next

Some complex tasks take longer than we would like, especially multi-step sequences that visit several pages in a row. We have a new navigation architecture in testing that looks like it will cut latency substantially, but I do not want to put numbers on it until the full 643-task suite finishes running against it.

The benchmark harness itself runs all tasks against real Chrome with Puppeteer, supports headless and headed mode, and can resume interrupted runs without losing progress. We run it in CI on every major architecture change.

Try breaking it

If you want to throw your own tasks at dassi, grab it from the Chrome Web Store. Plug in your own API key for Claude, GPT, Gemini, whatever model you prefer. And if you find a workflow where it falls apart, I actually want to hear about that more than I want to hear it worked, because the edge cases are where the next five percentage points live.