Is Training AI on Copyrighted Books Legal? Fair Use, Lawsuits and Authors Rights Explained

TL;DR
- In a pair of landmark rulings in June 2025, U.S. federal judges found that training AI on legally acquired books can be fair use, but training on pirated copies from shadow libraries like LibGen and Books3 is not, creating a major legal distinction for AI companies.
- Major lawsuits from authors including George R.R. Martin, Sarah Silverman, and The New York Times against OpenAI, Meta, and Anthropic are still moving through the courts, with outcomes that will define whether AI companies need to license content or pay damages.
- Authors and publishers argue unlicensed training threatens creative livelihoods and devalues human work, while regulators in the U.S. and EU are pushing for new transparency and licensing rules that could force AI firms to disclose training data and compensate creators.
The Battle Over Books and Bots
For years, AI companies quietly built their most powerful models by ingesting millions of books - from bestsellers and literary classics to niche nonfiction - often without asking permission or paying a cent. Now that practice is at the center of one of the most consequential copyright fights in decades. At stake is not just billions of dollars in potential damages, but the fundamental question of who owns knowledge in the age of artificial intelligence.
The controversy exploded when authors discovered their works had been swept up in massive datasets like Books3, which contained over 190,000 pirated books, and LibGen, used to train models including Meta's Llama and Anthropic's Claude. What tech companies call innovation, authors call theft on an industrial scale.
Why Fair Use Is the Heart of the Fight
In the United States, AI companies are betting their future on a single legal doctrine: fair use. The 1976 Copyright Act allows limited use of copyrighted material without permission for purposes like criticism, research, and commentary, judged by four factors: the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the market.
AI developers argue training is transformative. They aren't republishing books, they say - they're teaching an algorithm to learn patterns, grammar, facts, and reasoning from text, much like a human reading a library. The resulting model doesn't contain the books themselves, but a statistical understanding of language.
Authors and publishers counter that this stretches fair use beyond recognition. They argue the use is highly commercial, involves copying entire works verbatim, and creates products that directly compete with human authors by generating stories, summaries, and articles that could replace paid work. The U.S. Copyright Office weighed in with a major report in May 2025, warning that while some AI training might qualify as fair use, many instances - especially those that produce outputs substantially similar to the original works - go too far and require licensing.
Two Rulings That Changed Everything
Until 2025, no U.S. judge had directly ruled on whether AI training was fair use. That changed in late June with two back-to-back decisions that sent shockwaves through both Silicon Valley and the publishing world.
In Bartz v. Anthropic, Judge William Alsup in the Northern District of California ruled that Anthropic's training of its Claude models on books it had legally purchased and scanned was fair use, calling it "quintessentially transformative." However, he drew a hard line on piracy, ruling that Anthropic's use of more than seven million pirated books from LibGen and Pirate Library Mirror was not fair use and infringed copyright. The case will now go to trial to determine damages for the pirated copies, which could reach hundreds of millions of dollars.
Just two days later, in Kadrey v. Meta, Judge Vince Chhabria ruled in favor of Meta, finding that the authors who sued over the use of their books to train Llama had not presented sufficient evidence of market harm. But unlike the Anthropic decision, Chhabria emphasized his ruling was narrow and fact-specific, stating that other plaintiffs with stronger evidence could win. He also criticized Meta for allegedly using pirated datasets, calling it a separate issue that raised serious concerns.
Together, the rulings established a critical new legal framework: how you get the books matters as much as how you use them. Buying a book to train on might be legal; torrenting it is not.
The Wave of Lawsuits Still Ahead
Those two decisions are just the beginning. The most high-profile cases are still pending.
The Authors Guild and 17 prominent authors, including George R.R. Martin, John Grisham, and Jodi Picoult, sued OpenAI in late 2023, alleging ChatGPT was trained on their books and can generate near-verbatim summaries and passages. A similar suit by Sarah Silverman and other authors against OpenAI and Meta is also proceeding. Both have survived initial motions to dismiss and are moving toward discovery.
The New York Times' lawsuit against OpenAI and Microsoft, filed in December 2023, remains the most closely watched. While focused on journalism, it raises the same core questions about training and output, and the Times has presented examples of ChatGPT reproducing its articles word-for-word.
On the other side, Thomson Reuters scored an early victory for rights holders in February 2025 when a Delaware judge ruled that legal AI startup ROSS Intelligence's use of Thomson Reuters' headnotes to train a competing legal AI was not fair use. Though not about books, the decision is being cited by authors as precedent that training a direct competitor is not transformative.
Why Authors Say Their Livelihoods Are at Risk
Beyond the courtroom, authors say the practice poses an existential threat. The Authors Guild has repeatedly warned that generative AI can flood the market with cheap, AI-generated books that mimic human authors' style and compete for the same readers and royalties. Several authors have already found AI-generated knockoffs of their books for sale on Amazon.
The economic model of publishing is also under pressure. If AI companies can ingest entire libraries for free, authors argue, it devalues the years of labor that go into writing a book and removes any incentive for publishers to pay advances. Licensing, they say, is the only fair path forward - similar to how Spotify pays musicians or how publishers license books for audiobooks and translations.
Some publishers have already started to make deals. HarperCollins, Wiley, and Taylor & Francis have confirmed licensing agreements with AI companies, while The Associated Press and The Atlantic have struck content deals with OpenAI. For many authors, collective licensing is seen as the ideal solution: AI companies pay a fee to access a licensed catalog, with proceeds distributed to creators.
How Regulation Is Trying to Catch Up
Courts aren't the only arena where the future is being decided. Regulators on both sides of the Atlantic are racing to create rules for AI training.
In the U.S., Congress has held multiple hearings on AI and copyright, but no federal AI law has passed. The Copyright Office has recommended requiring AI companies to be more transparent about what works they train on and has launched an initiative to explore a voluntary licensing marketplace.
In the European Union, the AI Act, which entered into force in 2024, is now being implemented. It will require providers of general-purpose AI models to disclose summaries of their training data and comply with EU copyright law, including respecting opt-outs from rights holders. France and Germany are pushing for even stricter enforcement, requiring AI firms to prove they have obtained permission before training.
Meanwhile, AI companies are quietly changing tactics. OpenAI, Google, and Anthropic have all introduced opt-out tools for publishers and are increasingly emphasizing partnerships and licensed data in their newer models, a tacit acknowledgment that the era of scraping everything without consequence may be ending.
What Happens Next
The legal landscape is still deeply unsettled. The Anthropic and Meta rulings will almost certainly be appealed, and the OpenAI cases could take another year or more to reach trial. A Supreme Court showdown over AI and fair use is widely expected by legal experts within the next two to three years.
For now, the message to AI companies is mixed but clearer than ever before: training may be defensible as fair use, but piracy is not, and the burden of proving no market harm is on the tech firms. For authors, the twin rulings were both a setback and a validation - proof that courts are taking their claims seriously and that pirated training data carries massive legal risk.
The final outcome will shape not just the future of publishing, but what kind of internet we have - one where creative work is treated as raw material for machines, or one where human creativity remains a right that even the most powerful AI must respect and pay for.
Get All The Latest Updates Delivered Straight To Your Inbox For Free!