TechCrunch ran a piece with the headline “OpenAI reportedly finds evidence that more of its agents ran amok,” and I read it twice, mostly because “ran amok” is doing an enormous amount of work in that sentence. What I wanted from the story wasn’t which agent misbehaved. I wanted to know where it was running when it did, because that’s the variable that decides whether a browser agent going sideways is a shrug or an incident.

That detail almost never makes it into the reporting.

The model isn’t the interesting failure

Agents doing unexpected things is not news. These are probabilistic systems steered by natural language, and natural language is mush. Some percentage of the time an agent will misread an instruction, pick the wrong button, or decide that “clean up my inbox” includes the thread you were saving.

So the failure mode I care about isn’t the wrong click. It’s the wrong click that nobody saw for four hours.

Nobody has ever watched a progress spinner

The standard cloud agent interaction goes: you type a task, you get a spinner, you go make coffee, and eventually you get a summary written by the same system whose behavior you’re trying to evaluate. The intermediate steps exist somewhere, in a trace viewer or an execution log, and you could go read them if you were the kind of person who reads execution logs, which you are not.

Compare that to sitting next to a new hire while they do the same task. You’d catch the mistake at step three. Not because you’re vigilant, but because it’s happening in front of your face and wrong things look wrong.

Remote execution deletes that. The agent works on a machine in a datacenter, in a browser session you never see, and the only artifact you get is a report that says the task completed.

Watching is boring, which is exactly the point

Dassi runs in the Chrome side panel, so the agent works the tab you’re already looking at. The page scrolls. Fields fill in, character by character. Dropdowns open. When it clicks into the wrong Gmail thread, you see it click into the wrong Gmail thread, and you deal with it in the next two seconds instead of finding out in a summary later.

It is not glamorous. Most of the time you glance over, confirm it’s doing the boring thing you asked, and go back to what you were doing. But the review step costs you nothing because it’s ambient, and I’d take ambient oversight over an audit log every damn time.

We wrote about the visibility gap in Cloud Browser Agents Can’t See Your Tabs, mostly from the context angle. The oversight angle is the same architecture wearing a different hat.

Stop means stop

Close the panel. The agent stops. There’s no queued job finishing on a server somewhere, no half-committed action landing after you’ve walked away.

The credential problem underneath all of this

There’s a reason cloud agents need so much trust, and it isn’t ambition. An agent running on a remote server has no session state. It isn’t logged into your Gmail, your Salesforce, your bank, your team’s Notion. So to do anything useful on your behalf it has to be handed something: OAuth tokens with broad scopes, stored passwords, a remote browser profile you authenticated into once and then forgot about. Every one of those is a standing grant that outlives the task, keeps working after you close the laptop, and does its work in a place where you have no ambient view of what it touched.

Dassi doesn’t get a grant. It operates the session that’s already open in your browser, with whatever you’re currently logged into, and it loses that access the moment you log out. The permission and the visibility are the same thing, which is the part I find genuinely nice about the design rather than merely convenient. Martin Fowler’s write-up on agentic email called this the lethal trifecta: untrusted content, sensitive data, external communication. Running in a tab you’re watching doesn’t break the trifecta, but it does mean the third leg happens at human speed, in your field of view, with a close button.

More on the login-state piece in Your AI Browser Agent Can’t See That You’re Logged In.

What I’d want to know from OpenAI

Not “which agents ran amok.” That’s a headline. I’d want to know how long each one ran before anyone noticed, and what the person supervising it was looking at during that window. My guess is a spinner.

If you want the version where the agent works in front of you, Dassi is free on the Chrome Web Store, and you can bring your own key or sign in with the ChatGPT subscription you’re already paying for.

Watch it for the first few tasks. You’ll either trust it more or stop it faster, and both of those are wins.