Earnings
Home›Earnings›Previews›Anthropic piracy settlement reshapes AI training data…
Anthropic piracy settlement reshapes AI training data valuation
The settlement covers 482,460 works and is expected to pay about $3,000 per work, which is already prompting new licensing deals and tighter crawler rules for AI training.
A judge’s rulings in Anthropic’s piracy case are beginning to reshape how AI firms and investors think about the value of training data, as the dispute moves from individual platforms to broader licensing and consent standards. The settlement is expected to cover 482,460 works and pay roughly $3,000 per work, a narrower set than the more than 7 million pirated books identified in Anthropic’s broader library.
Mode Mobile CEO Dan Novaes said the decision is a turning point for “how markets value human data,” arguing that the human data market could expand into a $1 trillion opportunity by 2033. He linked the shift to legal pressure in the United States and Germany, saying stricter consent expectations are now reshaping how platforms monetize user activity and how data is used to train models.
The practical effect is showing up in new licensing and operational changes. Reddit has reportedly licensed its user-generated content to Google for roughly $60 million a year, while Cloudflare is set to begin blocking mixed-use AI crawlers from ad-supported pages by default starting Sept. 15, pushing AI labs to separate web indexing from model training and agent traffic.
Novaes also highlighted how companies may need more specialized human input as language models still hallucinate. He pointed to reinforcement learning from human feedback as a way to refine responses, and cited growth at Mercor, which he said rose from $1 billion to $2 billion in annualized revenue over four months this year, along with Meta’s 49% non-voting stake in Scale AI at a $29 billion valuation last year.