A new class action filed against Google adds another name to a growing list of AI providers accused of training their models on content they had no right to use. Publishers including Hachette, Cengage and Elsevier, along with author Scott Turow and the writers' group S.C.R.I.B.E., allege that Google trained its Gemini models on copyrighted books submitted under narrower, scope-limited programmes, Google Books and titles uploaded to the Google Play store, without ever securing permission to use them for AI training. The complaint goes further than a simple licensing dispute: it alleges Google altered or stripped copyright information from these works specifically to conceal that Gemini had been trained on material it was not authorised to use.
This is not an isolated incident. It follows a wave of similar disputes against Meta, OpenAI and Anthropic, and it arrives soon after Anthropic agreed to pay $1.5 billion, the largest payout in the history of U.S. copyright law, over its own use of pirated books in training. Two early California rulings found that training on copyrighted material can qualify as fair use, but that legal question is far from settled, and this new case has been filed in a different court, giving another judge the chance to weigh in. An internal Google document cited in the complaint reportedly acknowledged the practice could be "highly problematic" and carry fines running into the tens or hundreds of billions of dollars, evidence, the plaintiffs argue, that Google understood the risk and proceeded regardless.
This pattern is precisely the exposure Aphelion AI is built to remove. Aphelion AI is a private enterprise AI platform built to deploy your AI agent on your own infrastructure and your own data, rather than a shared foundation model whose training history is currently the subject of billion-dollar litigation. A business does not need to guess whether the intelligence layer it depends on was built honestly, because the entire question of stolen training data simply does not arise.
A Pattern, Not an Accident
What makes the Google case notable is not that a lawsuit exists, disputes over AI training data have become common, but the specific allegation that copyright information was deliberately altered to hide where the material came from. That is a different order of claim than an argument over what counts as fair use. It is an allegation of concealment, and it lands alongside Anthropic's own record-setting settlement for pirated training material as further evidence that some of the industry's largest players have, at various points, built their models on data they knew they should not have used, then dealt with the consequences after the fact rather than before.
- Google: accused of training Gemini on copyrighted books submitted for narrower purposes, and of concealing the practice.
- Anthropic: fined $1.5 billion for training on pirated books, the largest copyright payout on record in the U.S.
- Meta and OpenAI: facing their own pending disputes over unlicensed training material from authors and publishers.
Whatever the eventual legal outcome in any single case, the pattern itself is the point. A business that builds critical workflows on top of a hosted model has no visibility into how that model was trained, and as this year has shown, no guarantee that the answer would hold up in court.
$1.5 billion is the largest copyright payout in U.S. legal history, and it was paid by an AI lab over training data, not a hosted customer. That figure, and the "$10Bs to $100Bs" exposure Google's own internal document reportedly flagged, is the scale of risk sitting beneath any business relying on a shared foundation model whose training history it cannot verify.
Why Aphelion Does Not Carry This Risk
Aphelion's approach removes this exposure at the root rather than managing it after the fact. The platform does not scrape the open web or ingest unlicensed third-party content to build a foundation model of its own. Instead, it deploys your AI agent inside infrastructure you own or exclusively control, working from your own governed data, whether that is drawn through data enrichment of documents you already hold the rights to, or connected through system integration to CRMs, ERPs and document stores that are already yours. There is no separate, opaque training pipeline whose legality a customer is left to trust.
That distinction matters commercially as well as ethically. A business that depends on a hosted AI provider currently defending, or settling, a training-data lawsuit inherits a form of risk it did not sign up for: service disruption if an injunction is granted, sudden new licensing costs passed down through pricing, or simple reputational association with a company found to have concealed how its model was built. None of that risk transfers onto a business running a private deployment, because the platform underneath its operations was never built on data it had no right to use.
You should not have to trust that your AI provider was honest about where its model's intelligence came from. A private deployment means you do not have to.
What This Means for Due Diligence
For any business evaluating an AI vendor, this string of lawsuits is a reminder to ask a question that is easy to skip in a sales process: what data trained this model, and under what authority? Few hosted providers can answer that with full confidence right now, and fewer still can promise the answer will not change as more litigation works through the courts. Aphelion's AI agent sidesteps the question entirely, since it is built on your own data rather than a shared model with an unresolved legal history, and you can learn more about the people behind that approach on our About page.
None of this is an argument that AI cannot be trusted, it is an argument for how a business should access it. The safest position in an industry where the largest labs are being sued, fined, and accused of concealment over their training data is not to bet on any one of them clearing its name. It is to run AI on infrastructure and data you actually own, where the question of stolen content never gets a chance to become your problem.
Frequently Asked Questions
What does the new lawsuit against Google allege about AI training data?
A group of publishers and authors, including Hachette, Cengage, Elsevier and novelist Scott Turow, filed a class action against Google alleging that Gemini was trained on copyrighted books without permission. The suit claims Google took works submitted for narrower, scope-limited programmes such as Google Books and Google Play, and used them for AI training anyway, while allegedly altering or removing copyright information to obscure where the material came from.
How is Aphelion AI different from AI providers accused of training on stolen or unlicensed data?
Aphelion AI does not build or train a foundation model on scraped or unlicensed third-party content at all. It deploys a private AI agent inside your own infrastructure, working from your own business data and documents rather than a shared model whose training history is a legal question mark. There is no equivalent exposure to a lawsuit about where the underlying training data came from, because Aphelion is not in the business of quietly harvesting content to train on.
Does using a private AI deployment reduce a business's exposure to third-party AI copyright litigation?
Yes. When a business relies on a hosted AI provider whose training practices are under active legal challenge, that legal exposure can ripple outward through service disruption, sudden licensing costs, or reputational association with the case. A private deployment keeps your operations independent of any single provider's litigation history, since the platform your business runs on is not the same entity being sued over how it built its model.
Can Aphelion AI work with a business's own licensed content and documents without copyright risk?
Yes. Aphelion's data enrichment and system integration capabilities are built to work with the content your business already owns or is licensed to use, such as internal policy documents, product catalogues and CRM records, rather than material scraped from the open web. Because the platform operates on your own governed data inside infrastructure you control, you know exactly what went into the system and on what basis.
Private AI vs hosted AI providers facing copyright lawsuits: which carries more legal risk for a business?
A hosted provider facing an active or potential copyright suit over its training data carries a form of legal risk a customer cannot see, price, or control, and that risk sits underneath every workflow built on top of it. A private deployment built on a business's own licensed and owned data does not inherit that exposure, because there is no shared foundation model with a disputed training history standing between the business and its AI agent.