Anthropic's Own AI Broke Into Three Companies. Look at What It Was Handed.
TechCrunch ran it this morning: “Anthropic says its own AI models breached three companies during security tests.” Authorized exercises, consenting targets, nobody actually robbed. And my first instinct wasn’t to ask how smart the model was. It was to ask what somebody handed it before the clock started.
Because that part never makes the headline.
What the agent was given
Red-team exercises like this only work if the agent can reach real infrastructure. You can’t measure whether something moves laterally through a network if it never touches the network. So the setup includes credentials, remote access, and a leash long enough that no human is approving each individual step.
Which turns a story about capability into a story about provisioning. Anthropic’s models are good. They were also given keys.
I haven’t read a full technical writeup, only the coverage, so I’m reasoning from the shape these exercises usually take rather than the specifics of these three companies. That gap might change my read later. It hasn’t yet.
Standing credentials are the actual weapon
A model with no access is expensive autocomplete. A model holding a long-lived token to your production environment is a colleague who never sleeps, never checks in with anyone, and can be talked into things by a comment left in a Jira ticket by someone who doesn’t work at your company.
The second thing is what got breached. Not “AI.” A service account.
And that’s the pattern in nearly every agent deployment people are rushing into right now: spin up a VPS, mint an OAuth token with broad scopes, drop in an SSH key and a database URL, walk away. The token outlives the task. It outlives the agent. It’s still sitting there when the vendor holding it gets popped, which is a sentence I’ve now written about LiteLLM and several others since.
Nobody audits the leftovers. That’s the damn problem.
A browser agent can only reach the room you’re standing in
Dassi is a Chrome extension. It lives in the side panel, reads the page you currently have open, and clicks things in the tab you’re already looking at. There’s no server-side copy of your Gmail session, no OAuth grant parked in a third-party dashboard, no API key for your CRM stored in someone else’s cloud waiting to be inherited by whoever compromises them next.
The scope is your current browsing session. Nothing more. If you’re logged into Salesforce, the agent works in Salesforce; if you log out, it can’t. Close Chrome and the entire thing evaporates, because the session state never lived anywhere except your own machine. There is nothing to revoke afterward, since nothing was ever granted.
That’s a far less impressive security architecture than a sandboxed autonomous agent with its own identity in your org chart, and it produces a far smaller blast radius, which I’ll take every time. The worst case for a browser agent on my laptop is roughly “whatever I could have done in my browser in the next ten minutes, while sitting right there watching it happen.” The worst case for a credentialed autonomous agent is everything that identity was scoped to, for as long as the token stays valid, regardless of whether a single human is paying attention.
Visibility is doing real work here too. I can see the tab. I can see the click. If it starts opening the wrong record, I close the panel. That feedback loop doesn’t exist when the agent is running on a box in us-east-1 and reporting back in a summary written after the fact. We’ve written before about the three ways agents connect to your browser, and the differences look academic until a story like this one lands.
Where the local version still bites you
It’s not harmless. A browser agent operating in your already-authenticated tabs is operating on your real accounts, with your real permissions, and it can absolutely click send on a half-drafted email, archive the wrong thread, or submit a form with a number in the wrong column.
Prompt injection works here too. A malicious page can try to steer an agent that’s reading it, and the agent is reading whatever you put in front of it.
So the claim is narrower than “safe.” It’s that the mistakes are bounded by what you personally can do right now, they happen in front of you, and they end when you close the tab. Recovery means undoing an action, not rotating a credential and grepping logs for three weeks trying to figure out what else touched it.
The scope question is the only one worth asking
Every time someone pitches me on an autonomous agent, the demo is about what it can accomplish. Almost never about what it can reach. But reach is the number that decides how bad your worst day gets, and Anthropic just published three data points on that.
If you want the small-scope version, Dassi is free and lives in your existing Chrome profile: Chrome Web Store. Bring your own key or log in with your ChatGPT subscription.
Or skip it and go inventory your service accounts instead. Better use of an afternoon than reading one more post about how capable the models are getting.