PrismML spent today arguing that a tiny local model can cover most of what people actually use AI for, and the replies split about how you’d predict: half applause, half people pasting benchmark tables with Opus at the top. Both camps are answering a question I don’t think is the interesting one. Whether a small language model is “enough” depends almost entirely on how much the model is being asked to guess, and nobody trending today is measuring that.

Guessing is the expensive part. Parameters are what you buy to do it well.

What you’re actually paying frontier prices for

Point a cloud browser agent at a task. It opens a fresh Chromium instance somewhere in us-east-1, and that instance has never met you. No cookies, no session, no open tabs, no idea which of your four Google accounts is the one that matters. It hits a login wall and starts reasoning: this looks like SSO, that button probably goes to Okta, the record is probably under Opportunities and not Accounts, the export link is probably behind that gear icon.

Every single one of those “probably”s is inference. Inference is what large models are good at and small models are bad at. So when a cloud agent needs 400 billion parameters to book a meeting, a big chunk of that capacity is going toward reconstructing a context that was sitting on your screen the whole time.

The same task, two completely different jobs

Take something boring. You’ve got a Jira board open, forty tickets in the sprint, and you want the ones tagged for the mobile release summarized into a paragraph you can paste into Slack.

A cloud agent starts from nothing. It has to find your Jira instance, get past auth, figure out which board, figure out that your team spells the label mobile-rel and not mobile, work out that the sprint picker is the dropdown in the upper right and not the one that looks like a sprint picker but filters by epic, and then read a JQL-driven table that renders half its rows on scroll. That’s a planning problem with a dozen branches, and planning problems are where small models fall apart, because a wrong guess at step three poisons everything after it and the model has no way to notice.

An agent living in the tab you already have open starts somewhere else entirely. The board is rendered. The label filter you set last Tuesday is still applied. The DOM contains the ticket titles, the assignees, the status column, all of it as text, already scoped to exactly what you were looking at. The task collapses from “navigate an unfamiliar enterprise app while logged out” down to “read this structured text and write four sentences.”

The second job is not a frontier-model job. A 3B model does that. Honestly a decent 1B model does that, and it does it in 200 milliseconds on your laptop instead of two seconds in someone’s datacenter.

Context substitutes for parameters

That’s the whole argument. Every piece of real state your agent can see is reasoning it doesn’t have to perform, and reasoning it doesn’t perform is capacity it doesn’t need.

So use a cheap model for the boring 80%

This is where BYOK stops being a privacy talking point and starts being a budget one. Dassi runs in a Chrome side panel against the tabs you’re already authenticated in, and you pick the model per your own key. So you can point it at something small and cheap for the triage work: summarize this page, pull these rows, label this inbox, draft the two-line reply. Then switch to GPT-5.2 or Opus for the thing that needs an actual brain, like untangling a multi-step form flow across three subdomains. Swapping models in Dassi is a dropdown, not a migration.

Nobody sells you that split. Every closed AI browser routes you through one model at one price, and that price gets set by their hardest use case, not your easiest one. We wrote about why BYOK matters when providers keep changing the deal, and the cost argument turns out to be the same argument wearing different clothes.

Where the small model still faceplants

Ambiguity. If your prompt is vague, a small model picks the first plausible reading and commits, where a bigger one tends to ask or hedge. Long chains break too: anything past five or six dependent steps and the little ones start losing the thread of what they were originally doing.

And some pages are just hostile. Canvas-rendered dashboards, tables in iframes with no accessible text, apps that put everything behind <div> soup with class names like css-1x7hj9k. When the DOM is garbage, context stops rescuing you and you’re back to needing a model that can squint at a screenshot and reason. Cloud agents hit that same wall, except they hit it on every page instead of the bad ones.

The bet PrismML is making

They’re betting the ceiling on useful work is lower than the industry priced it at. I think they’re right, but only for agents that can see something. A tiny model in a chat box with no access to your world is just a worse chat box, and that’s most of what shipped this year.

The version worth building is the one where the model is small because it doesn’t have to be clever. It’s reading, not deducing. Give it the tab and it mostly just needs to not screw up the summary.

Which, for a 1B model in 2026, is a much lower bar than it used to be.