We ran dassi on Qwen3.8 27B, an open-weight model, on two benchmarks: Browser Use’s BU Bench V2 and Odysseys. It works, it’s cheaper than Gemini, and it is slow.

On BU Bench V2, dassi on Qwen3.8 27B scores 49.7 at $0.71 a task, against 17.8 for the same model in Browser Use's own table, with dassi on DeepSeek V4.1 Flash at 79.0 and on Gemini 3.8 Flash at 56.9

BU Bench V2: 49.7

On Browser Use’s public 55-task subset, dassi on Qwen3.8 27B scored 49.7, with 10 tasks meeting every rubric item. Same tasks, same judge and same reasoning setting as our DeepSeek and Gemini runs:

  • DeepSeek V4.1 Flash: 79.0, $0.08 a task, 14.7 minutes
  • Gemini 3.8 Flash: 56.9, $1.01 a task, 8.7 minutes
  • Qwen3.8 27B: 49.7, $0.71 a task, 43.9 minutes

Browser Use’s own table has this model at 17.8 and $0.69 a task. I wouldn’t read that as a head-to-head. Theirs is a different harness on a legacy 60-task cut, graded by a different judge. But the gap is large enough that it’s worth saying what we think it shows: for a mid-sized model, the agent around it matters a lot.

Odysseys: 27 of 40

On a 40-task sample of Odysseys (every fifth task), 27 tasks were fully completed (67.5%), with a mean rubric score of 74.4%. By difficulty that’s 7 of 9 easy, 8 of 10 medium and 12 of 21 hard tasks. Grading is the Odysseys authors’ own scorer, unmodified.

It’s a sample of 40, so there’s about 8 points of noise either way, and it isn’t directly comparable to the 85.0% we published for Gemini 3.8 Flash over all 200 tasks. It says Qwen is competitive here, not that it beats Gemini.

The catch: it’s slow

Qwen took 44 minutes a task on BU Bench V2, with a median of 83 model calls. Gemini took under 9. Our usual limit is 30 minutes, and in a pilot most tasks hit it, so we raised the limit to 60 for this run. Even then, 22 of the 55 BU Bench tasks hit it and were graded on what the agent had done by then. Another 8 errored out.

Some of that is the model, which generates slowly on the cheap hosts. The cheapest hosts on OpenRouter ran at 15 to 45 tokens a second, so we pinned the route to the two fastest cheap ones, at roughly 90 to 160. Some may be our own setup, covered below. Either way, a score of 49.7 undersells what a faster setup would do, and I can’t tell you by how much.

Cost

About $0.71 a task on OpenRouter, measured on our last 15 tasks, which were a clean run with no failed work. That’s near Browser Use’s $0.69 for the same model. It is also nine times what DeepSeek V4.1 Flash cost on the same benchmark.

The whole effort cost $110.55, for two benchmarks. A good part of that, which I’d estimate at about a quarter, bought nothing, because of what happened next.

What went wrong

We meant to finish in 36 hours and didn’t. Four of our CI shards were killed by the kernel for running out of memory, and each one took every finished task in it with it, since a killed run uploads nothing.

We read the host’s kernel logs to find out why. Each CI container has a 4 GiB memory limit, and Chrome filled it: 29 to 89 Chrome processes holding about 4 GB, while dassi’s own extension used under 600 MB. We then reproduced it locally on one heavy Odysseys task. One ad-heavy news page spawned about 700 iframes, mostly ad and cookie-sync frames, and Chrome turned the run’s iframes into roughly 120 renderer processes. We tried a cap on how many tabs the agent can open, but the agent never reached it in our test, so we learned only that tabs weren’t the whole story. The ad frames were the bigger part. The slowness may partly be that memory pressure, since a container thrashing against its limit is slow, but we haven’t proven that.

What did work was running small shards of three or four tasks, so a kill loses a few tasks instead of twenty. Six BU Bench tasks and nine Odysseys tasks were re-run once after their first attempt was lost, and the re-run is the one that counts.

How far to trust this

  • Both runs used dassi 0.89.0 on one attempt per task, at the high reasoning setting, with a 60-minute limit where the other runs above used 30. We also pinned the OpenRouter route to two hosts for this run. That pin isn’t in any dassi release.
  • BU Bench V2 is graded with Browser Use’s own judge code on Gemini 3.8 Flash, not the GPT-5.6 Luna they use. Odysseys uses the authors’ scorer on its default judge model. Judge noise is real, so only the means over many tasks mean much.
  • None of these scores has been checked by either benchmark’s maintainers.

Every task’s id, score, status, token counts and time are in the evals repo, with no task text or traces, as Browser Use asks.