dassi on Gemini 3.8 Flash Scores 57 on BU Bench V2, and a Model 12 Times Cheaper Beats It
Our first post about Browser Use’s BU Bench V2 said dassi on Gemini 3.8 Flash scored 84.7. That number was wrong, and we’ve taken the post down.
It came from our own copy of the benchmark’s judge, running on a small model (Gemini 3.5 Flash-Lite), and that judge waved nearly everything through. When we later ran Browser Use’s actual judge code, the same kind of answers that ours had marked 100 came back at 88, 48 and 0. So we reran Gemini on Browser Use’s public 55-task subset and graded it properly, next to the DeepSeek V4.1 Flash run. Same agent, same tasks, same judge, same reasoning setting.

Gemini scored 56.9. DeepSeek scored 79.0, for a twelfth of the money.
The two runs, side by side
- Score (0 to 100), 55 tasks: Gemini 56.9, DeepSeek 79.0
- Tasks with every rubric item met: Gemini 5, DeepSeek 18
- Tasks zeroed for reward hacking: Gemini 5, DeepSeek 0
- Tasks under 50: Gemini 16, DeepSeek 4
- Model cost per task: Gemini $1.01, DeepSeek $0.08
- Minutes per task: Gemini 8.7, DeepSeek 14.7
Head to head, DeepSeek did better on 40 of the 54 tasks that ran. Gemini won 9, and 5 were ties. One task never started for either model because its website stopped resolving that day, and it counts as zero for both.
The judge here is Gemini 3.8 Flash itself, running Browser Use’s prompt and scoring code. If a model grading its own family has a soft spot, it helped Gemini, which makes the gap harder to explain away.
Fast, and a bit too willing to fill in blanks
Gemini finished tasks in under 9 minutes, against almost 15 for DeepSeek. That speed didn’t come for free.
Five Gemini runs were zeroed outright for reward hacking, which on this benchmark means the judge found deliverables the evidence can’t support. In one, the final report listed a news article the agent had never opened. In another, it built LinkedIn job links from made-up sequential IDs, and those pages turned out to be unrelated jobs. The other misses were more ordinary: a table cut off at 50 rows when there were 235, or a classification rule applied wrongly across a whole dataset.
None of DeepSeek’s runs was zeroed that way. It was slower, and I suspect that’s the same thing seen from the other side: more checking, fewer invented rows. I’d rather have that from an agent working in someone’s logged-in accounts, where a confident wrong answer is worse than a late right one.
Why Gemini costs twelve times as much
The two models read roughly the same amount per task, and almost all of it is cached conversation history.
- Cached input: Gemini 6.4M tokens for $0.48, DeepSeek 4.8M tokens for $0.014
- Uncached input: Gemini 375k tokens for $0.28, DeepSeek 179k tokens for $0.027
- Output: Gemini 68k tokens for $0.25, DeepSeek 65k tokens for $0.039
- Total per task: Gemini $1.01, DeepSeek $0.08
The biggest line is cached input. DeepSeek charges $0.003 per million cached tokens off-peak; Gemini 3.8 Flash charges $0.075, 25 times more. Output is about 6 times more on Gemini, and uncached input 5 times.
Gemini also has twice as many uncached tokens, and we tracked that down with a small experiment against its API. Gemini’s automatic cache only extends in steps of about 4,000 tokens, so every call re-bills the part of the history past the last step, a couple of thousand tokens on average. DeepSeek caches in 64-token units, including the model’s own previous reply, so it only pays for what’s actually new. That quirk adds roughly 10 cents per Gemini task. The rest is just the price list.
Two more things to keep in mind. The DeepSeek run fell entirely in its off-peak hours, and at peak its cost doubles to $0.16. And Gemini 3.8 Flash is on an introductory price until the end of the year; from January the same Gemini run would cost about $2 a task.
How far to trust this
Both runs used dassi 0.80.0 unmodified, at the high reasoning setting dassi uses, with one attempt per task and a 30-minute limit. The judge is Browser Use’s own code and rubrics, but on Gemini 3.8 Flash rather than the GPT-5.6 Luna they use, and their published chart is on an older 60-task cut. Browser Use asks people not to publish decrypted tasks or traces, so we haven’t, but every task’s score, rubric counts, tokens and cost for both models are in the evals repo, with the script that runs Browser Use’s judge on saved runs.
I went into this expecting Gemini to be the safe choice and DeepSeek the cheap experiment. On this benchmark, at least, it went the other way round.