You probably already know that the AI models behind ChatGPT, Gemini, Claude, and other chatbots are trained on a seemingly endless database of published works, including hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. Most published authors are contributing to the development of the same AI tools that threaten to harm their lives, without their knowledge or consent. That seems illegal, right?
The reality is not so simple.
“I think one of the problems with this whole legal field and this whole technology field is that there’s a lot going on,” Kathy Gellis, an attorney with expertise in intellectual property, copyright and technology, told TechCrunch. “It’s very complicated and there’s a lot of raw emotion, both for and against what’s going on.”
Last year, in one of the first rulings of its kind, Judge William Alsup ordered Anthropic to pay a massive $1.5 billion copyright settlement to a group of authors whose work was used to train the company’s AI models. At face value, this seemed like a moral victory in the author’s favor, but Judge Alsup actually ruled that Anthropic’s AI training was legal. Alsup punished Anthropic for pirating these books from an illegal online shadow library.
“Like readers who aspire to be writers, Anthropic’s LLMs trained their work to turn difficult corners and create something different, rather than rushing forward to imitate or supplant,” the judge wrote, likening the LLM’s ingestion of trillions of words to a writer’s literary studies.
Gellis believes the ruling is more favorable to AI companies. How much is a $1.5 billion fine for a company that projects annual sales of approximately $200 billion by 2028?
“I think it’s generally good news for training AI that he saw what was going on and actually thought it was akin to reading a copyrighted work rather than copying a copyrighted work,” Gellis said. “Copyright law depends on copying, but not on using, experiencing, consuming, or reading the work.”
Copyright law hasn’t been updated since 1976, so judges must consider how to interpret 50-year-old guidelines when facing legal questions that could shape the future of the AI industry.
“Everyone is very worried right now, because the laws are all over the place, and that’s because of this issue,” Jason Henderson, senior attorney and founder of the IP & Media practice at JWL International, told TechCrunch. “They know that AI models have been trained on so many things, but the law hasn’t quite caught up with the questions.”
These issues often hinge on fair use laws, that is, whether the use of a copyrighted work is sufficiently “transformative” to be legally permissible.
Fair use carves out copyright law that allows the use of copyrighted material without explicit permission, and protects the ability to comment on or repeat copyrighted works through criticism, parody, education, or other means. Judges consider certain factors when determining whether something is fair use, such as the purpose and nature of the work, the amount used, and the impact on the market.
“Copyright has always been about protecting and growing markets,” Henderson said. “Courts’ reasoning[in AI cases]is kind of disjointed. What tends to win is if you’re trying to compete directly, and you’re training on someone else’s property, courts will frown on it…If what you’re doing isn’t meant to be competitive, courts tend to find a way to make it okay.”
Henderson was referring to a lawsuit in which media technology company Thomson Reuters sued research firm Ross Intelligence for copying content to build a competing AI-based legal platform.
Judge Stefanos Vivas said last year that “Mr. Ross’s use is not transformative because it has no ‘additional purpose or different character’ from Mr. Thomson Reuters’s use.”
In this case, Judge Vivas ruled that training Reuters content to create a new platform that directly competes with Reuters was not fair use. Authors may argue that chatbots compete with them by using their copyrighted material to generate new composite books, but that claim has not yet been successfully argued in court.
When it comes to the relationship between AI and copyright, Gellis thinks it’s helpful to narrow down what we’re actually talking about. The way we think about copyright in terms of AI training is very different from the way we think about copyright protection for AI-generated content.
In one case, Saylor v. Perlmutter, a court ruled that a work is not copyrightable if it is 100% generated by AI, which opens up a whole new can of worms. How can you definitively prove whether a work was generated using AI? And if so, how do we know what percentage of that work is or was created by AI?
“If you write a novel in[Microsoft’s]Word and run spell check, we kind of resist the idea that Word doesn’t own your novel,” Gellis said. “[AI]is forcing us to consider a plethora of decisions that we have previously ignored.”
Most AI companies still have pending litigation over these issues, which means definitive solutions to these issues may not be forthcoming for some time.
“What you’re seeing is that the first salvo has an impact, and that impact itself could be undone if another court rules differently. We’ll need the status of the subsequent cases to determine who wins,” Gellis said. “But in the meantime, all these decisions are shaping everything that’s happening. It would be foolish in some ways for AI companies to ignore them.”
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
