Browser Use’s BU Bench V2 chart has one dot I kept coming back to: DeepSeek V4.1 Flash at high reasoning, 41.5 points for 15 cents a task. It’s a cheap model, and on their harness it lands about where cheap models usually land.

We put the same model inside dassi and ran Browser Use’s public 55-task subset. Graded with Browser Use’s own judge code, it scored 79.0, and a task cost 8 cents.

dassi on DeepSeek V4.1 Flash scores 79.0 for $0.08 a task on BU Bench V2, next to Browser Use's 40 published entries; their run of DeepSeek V4.1 Flash at high reasoning scored 41.5 for $0.15, and GPT-6 Astra scored 80.6 for $8.87

That puts it level with the top of Browser Use’s chart, GPT-6 Astra at 80.6, which spends $8.87 a task to get there. Roughly a hundred times the money for a point and a half.

First, a correction

An earlier post reported 84.7 for dassi on Gemini 3.8 Flash, graded by our own port of the judge running on Gemini 3.5 Flash-Lite. That judge was much too kind. When we later ran Browser Use’s actual judge code on three of those same Gemini tasks, the scores went from 100, 100 and 100 to 88, 48 and 0 (the zero was flagged for stating product facts from pages the agent never opened).

So we threw our judge out and took that post down. Every number here comes from Browser Use’s own grading code. We reran Gemini 3.8 Flash the same way and it scored 56.9, which has its own post.

What the 79 is made of

Of the 55 tasks, 18 met every rubric item and 23 scored 90 or better. Eleven more landed in the 80s, 16 between 50 and 79, and five under 50.

The misses are the interesting part. Four tasks ran into the 30-minute limit and were graded on whatever the agent had done by then, which came to 4, 24, 32 and 53. One task never started because its website stopped resolving the day we ran it, and it counts as a zero. Nothing was re-run.

For contrast, our old judge gave the same DeepSeek run 97.3 on the tasks it could grade. The gap between 97.3 and 81.4 (those same 53 tasks under Browser Use’s code) is basically the difference between a grader who reads the final answer and one who checks the answer against what the agent actually saw.

8 cents, measured

We added up every model call’s tokens and priced them at DeepSeek’s list rate.

  • Per task: $0.08 on average, $0.079 median, from $0.03 to $0.14.
  • The whole run: $4.34 for 54 tasks.
  • Why it’s that low: 96% of the input tokens were cache hits, which DeepSeek charges at a fiftieth of the normal input price. Output tokens were the biggest line item, $2.12 of the $4.34.
  • The fine print: every call happened to fall in DeepSeek’s off-peak hours. At peak prices the same run averages $0.16 a task. Off-peak it’s cheaper than every entry above 45 points on Browser Use’s chart; at peak, two GPT-6 Luna settings sneak under it, at scores of 50 and 54.

A task took 14.7 minutes on average and about 64 model calls. That’s slower than Gemini 3.8 Flash in dassi, which averaged under 9 minutes, but for long unattended errands I’ll take slow and cheap over fast and wrong.

Same model, same setting, 37 points apart

I don’t think the model is the story here. Browser Use ran DeepSeek V4.1 Flash too, at the same high reasoning level dassi uses, and got 41.5 (their best setting, max, reached 47.9). We got 79.0. The weights and the setting are identical, so almost all of that gap is the harness: how the agent sees the page, what it can do in one step, and whether it keeps going when a site throws a CAPTCHA or a half-rendered grid at it.

dassi reads pages as structure (elements, labels, state) rather than pixels, and a single model call can run a small program that works through a whole grid of product cards at once. That’s my best guess at where most of the 37 points come from, though we haven’t split it out cleanly, and I’d rather say that than pretend otherwise.

How far to trust this

Pretty far, with three asterisks.

The judge is Browser Use’s code but not Browser Use’s model. Their chart grades with GPT-5.6 Luna at xhigh reasoning; we ran the identical prompt, schema and scoring on Gemini 3.8 Flash. As a sanity check we also graded everything with Gemini 3.1 Pro, which gave 81.9, so the result doesn’t seem to hinge on which Gemini sits in the chair. Luna could still be stricter than either of them.

The task sets differ slightly. Browser Use’s chart uses an older 60-task cut. Their public 55-task subset is the overlap they recommend reporting on, and it’s the one we ran, at the exact dataset revision their subset file pins.

And we graded ourselves. Browser Use asks people not to publish decrypted tasks or traces, so we can’t post the sessions, but every task’s score, rubric counts, token counts and cost are in the evals repo, along with the script that runs Browser Use’s judge on saved runs. Anyone with the benchmark can check the grading.

Gemini 3.8 Flash has since had the same treatment. Next is the whole thing again with Browser Use’s actual judge model once we’ve wired up the key. If Luna knocks us down, we’ll say so here. I’m hoping it doesn’t, though I’ve been wrong about judges once already this week.