Opus 5 Is Here. It's Still Sitting in a Logged-Out Chat Window.
Anthropic shipped Opus 5 this morning and by lunch my feed had turned into the usual carnival: benchmark screenshots, three “this changes everything” threads, and one guy insisting it’s Opus 4.8 with a fresh coat of paint. I gave it an hour of real work. It’s better at long multi-step stuff, the kind where step seven depends on something it noticed back in step two and quietly held onto.
Then I asked it about an invoice sitting in my Gmail tab and remembered where I was. A chat window. Logged into nothing.
The model was never what was broken
For about two years the standing excuse for why AI assistants couldn’t do your actual job was model quality. Not smart enough yet. Wait for the next one. And each release did close some of that gap, Opus 5 included.
But go look at where your assistant actually failed you last week. Mine failed because it couldn’t read a Notion page behind SSO, couldn’t see the seven tabs I had open comparing vendor pricing, couldn’t tell that the Salesforce record I was staring at had a stale owner field, and none of those are reasoning problems — they’re location problems. A model that scores five points higher on SWE-bench fixes exactly none of them.
So the bottleneck moved. It’s not what the model knows. It’s what the model can see when you ask.
The remote-browser wave has the same hole in it
Hacker News had a good day today too. Several browser agent launches, at least two of them open source, most of them advertising autoscaling browser pools as the headline feature. They’re well built. They also all run the browser somewhere else.
Which means the browser boots clean. No cookies, no session, no history, no logged-in anything. The fixes are all worse than the problem: export your cookie jar and upload it, proxy your session through their infrastructure, or log in again inside a remote Chrome you don’t control and type your 2FA code into a box on someone else’s server. That last one is a hell of a thing to ask a person to do casually, and yet it’s in the quickstart docs of half these projects.
I wrote more about that pattern in why three autoscaling browser agents launched in one day and still solved the wrong problem. Nothing about Opus 5 changes that math. A smarter model in a logged-out browser is a smarter model that still hits a login wall.
Why bring-your-own-key matters specifically on launch day
Here is the thing BYOK quietly buys you, and it’s not privacy this time. It’s timing.
When a product hardcodes its model provider, a new release means waiting for a vendor integration, a pricing negotiation, a version bump, a changelog entry six weeks later. With a key you own, launch day is a dropdown. You paste the key, pick the model, and you’re running Opus 5 inside your own Chrome before the release blog post finishes loading.
That’s how it works in Dassi. It sits in the Chrome side panel, reads the tab you’re already looking at, and calls whatever model your key points to. Anthropic, OpenAI, Google, DeepSeek, whatever you’ve got. Or you sign in with a ChatGPT subscription and skip keys entirely. The session state is yours because it’s literally your browser profile, not a container in us-east-1 pretending to be you.
Same weights. Different room.
The unglamorous bit
None of this is a model capability. Opus 5 in a side panel and Opus 5 in a chat tab are byte-identical. The only difference is what’s in the DOM when the request goes out.
What I’d actually test with it
I’m going to spend this week running Opus 5 against the tasks that used to need babysitting: long Gmail threads where the ask is buried in message four, multi-tab research where the agent has to hold six pricing pages in its head, form-filling on admin panels built by people who hate you. Those are the ones where the extra reasoning depth should show up as fewer corrections from me, and if it doesn’t, I’ll say so.
We ran a similar test on why your AI agent can’t see that you’re logged in, and the failure modes were embarrassingly boring. Not hallucination. Just doors it couldn’t open.
Anyway. Go swap your model dropdown. It takes eleven seconds, and it’s more useful than reading another benchmark chart from a company that picked which benchmarks to show you.