Anthropic Built a Code Review Tool. For Its Own AI's Code.
Anthropic announced a code review tool last week designed specifically to catch problems in AI-generated code, and the framing was almost apologetically honest. The company that builds Claude — the model that thousands of developers use to write code every day — is now shipping a product whose entire purpose is to double-check Claude’s work. That should tell you something about where we actually are with AI autonomy, which is nowhere close to “let it run unsupervised.”
The company that makes the AI doesn’t trust the AI
I keep seeing people on Twitter talk about fully autonomous coding agents like they’re three months away. Meanwhile, Anthropic just invested engineering resources into building a review layer because the code its own model generates is not reliable enough to ship without human oversight. And Anthropic is not some cautious laggard — they’re one of the companies pushing hardest on agent capabilities.
The review tool scans for security vulnerabilities, logic errors, and the kind of subtle bugs that LLMs introduce because they’re pattern-matching on training data rather than reasoning about your specific codebase’s constraints. It catches things like improper input validation, race conditions that only manifest under load, and dependency conflicts that look fine in isolation but break when deployed alongside existing code. These aren’t edge cases. This is the bread-and-butter stuff that junior developers get wrong too, except an LLM generates it with complete confidence and no hesitation, which somehow makes it worse because you’re less likely to scrutinize something that reads like it was written by someone who knew what they were doing.
Code is just the canary
So code review for AI output is now a product category. But the principle underneath it extends way beyond programming.
Every time you let an AI draft an email, fill out a form, summarize a document, or pull data from a webpage, you’re in the same position as a developer staring at AI-generated code. The output looks plausible. It reads fluently. And it might be subtly wrong in ways that are hard to catch unless you’re actively watching.
I’ve seen this firsthand with browser tasks. An AI agent filling out a shipping form got the zip code right but transposed two digits in the street address because it pulled from a similarly-named contact. An email draft sounded perfect but referenced a meeting that happened last Tuesday when the actual meeting was last Thursday. Small errors, big consequences, and the kind of mistakes that only get caught if you’re looking at what the AI is doing in real time rather than reviewing a finished result after the fact.
Why the “run it in the background” pitch is wrong
The hot narrative in AI tooling right now is autonomy. Let the agent handle it. Go do something else. Come back when it’s done. And sure, that sounds efficient — until you remember that Anthropic, the company building these agents, just told you their output needs a second pair of eyes.
Background autonomy works for tasks with low stakes and easy reversibility. Renaming a batch of files, maybe. But the tasks where AI could save you the most time — email, research, form submissions, data entry across web apps — are exactly the tasks where mistakes have real downstream effects. You cannot un-send an email. You cannot un-submit a reimbursement claim with the wrong amount.
This is why I keep coming back to the side panel approach that dassi uses. The AI works right next to you, visible in your browser, and you see every action before it happens. It is not doing things behind your back while you’re in another tab hoping nothing goes wrong. You are the code review tool, except for browser tasks instead of code.
Birgitta Boeckeler’s “harness” concept fits here perfectly
Birgitta Boeckeler wrote recently about the idea of “harness engineering” — building constraints, guardrails, and verification layers around AI agents to keep them in check. The OpenAI team she referenced spent five months building these harnesses before they could trust their AI-generated codebase, and even then the focus was on deterministic checks, structural tests, and human-in-the-loop verification rather than just hoping the model got it right.
Browser tasks need the same philosophy. The harness for a browser agent is not a linter or a test suite. It is your eyes on the screen. It is the ability to see what the agent is about to do, approve it or correct it, and move on. That only works if the agent is doing its work where you can see it, not in some headless browser running on a server somewhere.
The trust gradient nobody talks about
There’s a spectrum between “AI does everything” and “AI does nothing” and almost every useful application of AI right now sits somewhere in the messy middle. Anthropic’s code review tool is an admission that the right position on that spectrum, for code, is “AI writes it and a separate system checks it.” For browser tasks, the right position is probably “AI proposes actions and you confirm them,” which is exactly what a side panel agent gives you.
The people building fully autonomous agents will get there eventually. But “eventually” is doing a hell of a lot of heavy lifting in that sentence, and in the meantime the most practical approach is AI that works with you rather than instead of you. Anthropic apparently agrees, even if they’d phrase it more diplomatically.