Publishers Sued OpenAI Again. Your Browser Reads What You Already Pay For.
The Seattle Times and Newsday both filed against OpenAI and Microsoft this week, which puts them somewhere around twentieth in line behind the New York Times, and the complaints all rhyme. Crawlers pulled the articles. The articles landed in training runs. Then they landed in retrieval systems that answer the question instead of sending the reader to the page.
What strikes me is that none of these cases are about reading. They’re about acquisition. Who fetched the copy, what server it came from, whether anyone had a subscription at the moment of the fetch, and where that copy sat afterward.
The fight is over the fetch
Strip the legalese out and every publisher complaint is a claim about provenance. A machine in a data center requested a page it had no subscription to, kept the text, and shipped it inside a product. The subscriber isn’t the defendant here. The pipeline is.
Nobody has sued anyone over the tab they had open
Because that’s just reading. You paid, you logged in, the server sent you the page. Whatever software renders it for you afterward is downstream of an access that already happened, lawfully, with your credentials.
What happens when I ask my sidebar to summarize something I pay for
Say I’ve got a Seattle Times investigation open. I subscribe. The paywall already checked my cookie, the server already decided I was allowed, and the full text is sitting in the DOM of a tab on my laptop.
When I ask Dassi to condense it into five bullets for a colleague, no new request goes out to seattletimes.com. There’s no crawler, no cache, no third-party fetch, no copy landing in an index that gets queried later by someone who never paid. The agent reads the rendered page, the same bytes my eyeballs are reading, and the model call goes to whichever provider’s key I brought. That’s the whole loop.
I want to be precise about the part that isn’t magic, because a lot of privacy marketing in this space is bullshit. The article text does leave my machine when the model is a hosted one. It goes to my Anthropic or OpenAI account, under my terms of service, as a prompt in a conversation I initiated. What doesn’t happen is a copy accumulating in someone’s retrieval corpus to be served back to strangers, which is the specific behavior these lawsuits describe. We wrote more about that routing distinction in AI Browser Safety: What Gets Sent Where, and it matters more than most people assume, because “local agent” and “your data never leaves” are claims that quietly stop being true the second a cloud model is involved.
The extension is on the Chrome Web Store if you want to poke at the behavior yourself.
The research portal case is messier and I’d rather say so
My company licenses a market research portal. Ugly interface, no export, no API on our tier. Pulling twelve data points out of it by hand takes twenty minutes and I get one of them wrong roughly every other time.
A browser agent reading that rendered table is doing something my license clearly permits at the access layer. I’m an authorized seat. But some enterprise licenses have clauses about automated extraction or systematic downloading that don’t care whether the credentials were valid, and if you’re at a company with a procurement team, someone should read the contract before you point an agent at the vendor’s dashboard. I’m not a lawyer, and the honest version of this section is that access rights and use rights are separate questions.
Still a different animal from an unauthorized crawler. Nobody at that vendor has to wonder how their content got into a product they never licensed.
”Your session did it” versus “our crawler did it”
Those are different sentences about different acts, and the entire publisher litigation wave turns on which one applies.
An agent operating inside your authenticated session is exercising an authorization you were granted and paid for. It inherits your permissions, your rate limits, your account’s terms. When it can’t get in, you couldn’t get in either. That property is why we keep writing about logged-in tabs as an architecture rather than a convenience, like in The API Needs Approval. Your Logged-In Tab Doesn’t.
Meanwhile the authors contesting the Anthropic settlement terms are arguing about a fund covering books that were acquired, in some cases, from pirate libraries. The through-line across all of it is sourcing. Not intelligence, not capability, not model quality. Where the copy came from.
Which is a slightly funny place for the industry to have arrived, given how much money went into pretending the answer didn’t matter.