← All writing
AI Jul 22, 2026 6 min read

A judge just put a price on where training data comes from

Christopher Dorsey

Christopher Dorsey

AI & MadTech Advisor · Enterprise Sales Leader

TL;DR

On July 21 Judge Araceli Martínez-Olguín gave final approval to Anthropic's $1.5 billion settlement with authors: about $3,000 per book across more than 482,000 pirated titles, over 91% claimed, the largest copyright settlement in US history and the first of the major AI training cases to reach the end. The line that survives came from Judge Alsup's earlier ruling: training on lawfully acquired books is fair use; acquiring them from piracy sites is infringement no matter what you built. Anthropic paid for the acquisition, not the training. The coverage is authors and checks. The part that lands on anyone selling AI: procurement now has a reference price for the training-data provenance question, and it will show up in security reviews next to SOC 2. Indemnification splits the field — OpenAI, Microsoft, Google and Anthropic offer copyright indemnities with real caps and conditions attached, mostly covering outputs rather than training data; the 40-person AI app in the same bake-off offers a promise it couldn't survive honoring. If you sell AI, write the provenance answer down and put your indemnity chain in the standard security packet. If you buy, ask whether the indemnity covers the training data or only the outputs, and read the cap out loud before you rely on it.

On July 21, Judge Araceli Martínez-Olguín gave final approval to Anthropic’s $1.5 billion settlement with the authors whose books were pirated to train Claude: roughly $3,000 per book across more than 482,000 titles, with over 91% of them already claimed and $101 million in attorneys’ fees riding along. It is the largest copyright settlement in US history and the first of the big AI training cases to reach the end, while suits against OpenAI, Meta and others grind on.

The legal line worth memorizing came earlier, from Judge William Alsup, before he retired: training a model on books you lawfully bought is fair use. Downloading them from piracy sites is infringement, and it stays infringement no matter how transformative the model that ate them. Anthropic paid $1.5 billion for how it acquired the data. The training itself was never the exposure.

The coverage this week is authors and checks, and fair enough; $3,000 a book is real money for a working novelist. The settlement also handed a number to your buyer’s procurement team, and that is where it starts changing deals.

I’ve watched a footnote become a deal gate before

At Zeta I sold data-driven marketing through the years when GDPR and CCPA gave the question “where did this data come from” legal teeth. Before that, provenance lived in a footnote of the MSA. After, deals passed through privacy reviews that hadn’t existed a year earlier, and vendors who couldn’t produce a clean sourcing answer stopped making shortlists. Nobody announced the change. The checklists just grew a row.

Training data is next, and this time the row arrives with a dollar figure attached. Statutory damages for willful copyright infringement run up to $150,000 per work, and Alsup’s framework put Anthropic’s theoretical exposure into numbers no board could sit with, which is why the case settled. Risk teams can’t price “maybe AI has a copyright problem.” They can price a settled class action with a per-book rate. The security questionnaire that already asks about agent governance is about to grow a provenance section.

Indemnity splits the field

The frontier labs saw this coming. OpenAI has offered its Copyright Shield to enterprise and API customers since 2023; Microsoft has the Copilot Copyright Commitment; Google indemnifies across much of its generative stack; Anthropic’s commercial terms commit it to defend customers over authorized use. Read the fine print, though, and the protection is narrower than the press releases: most of these cover what the model outputs, not the vendor’s own training-data liability, and the conditions and caps vary wildly from contract to contract.

An indemnity is also only as good as the balance sheet behind it. Anthropic could write a $1.5 billion check and keep operating. The 40-person AI app in the same bake-off cannot, and its indemnity clause, if one exists, is a promise from a company that would not survive the claim. If you sell an AI product built on somebody else’s model, you inherit the stack’s provenance whether you like it or not, and your buyer’s counsel will eventually map the chain: whose model, trained on what, indemnified by whom, up to how much. Map it first.

Get the answer in the packet

If you sell AI: write the provenance answer down before somebody makes you. Where the training data came from, what was licensed, what your model vendor indemnifies and to what cap, and what you cover on top. Put it in the standard security-review packet next to the SOC 2 report. The first vendor in a category to answer cleanly sets the bar the rest get graded against, and this question is cheap to answer today and expensive to answer under deadline in a stalled deal.

If you buy AI: two additions to the evaluation. Ask whether the indemnity covers the training data or only the outputs. Then read the cap out loud in the meeting, because a $500,000 cap against a $150,000-per-work exposure is a coupon, not coverage. And ask the provenance question even when the demo is beautiful. Anthropic had the best model in the world and still wrote the largest copyright check in American history for the way it filled its bookshelf.

Share this post

About the author

Christopher Dorsey

Christopher Dorsey

Enterprise Sales Leader · AI Go-To-Market · Startup Advisor · Denver, CO

Fifteen years selling technology to Fortune 500 brands across AI, advertising, and data infrastructure — most recently at Zeta Global, Oracle, and Fastly. Currently advising founders and sales leaders on AI go-to-market and Generative Engine Optimization.

Questions, pushback, or just want to compare notes?