Can AI train on copyrighted work? The government hopes so


The Justice Department urged a judge to side with OpenAI and Microsoft in a lawsuit brought by The New York Times, saying artificial intelligence companies should be allowed to train their models on material owned by the news giant.

It’s a move that could further smooth the way for future AI training practices, if a judge ultimately rules against the Times. The lawsuit, filed in 2023 in Manhattan Federal District Court, states that using the paper’s content to train chatbots breaches copyright protections and the fair use doctrine, a legal exception that allows the unlicensed use of copyrighted material in certain circumstances.

“Microsoft and OpenAI stole from The New York Times to make commercial products that substitute for its journalism, threaten its business, and undermine its industry,” The New York Times wrote in a Sept. 4 court filing.

As large language models have become more popular since ChatGPT’s release in 2022, questions over what they train on and who gets credited have been simmering.

Many authors, songwriters and other creative artists have engaged in legal action seeking compensation for their work being used for AI training, totaling over 100 current copyright lawsuits against AI companies, according to a comprehensive list kept by Santa Clara Law School professor Edward Lee.

The New York Times’ complaint asks for monetary damages, a court order preventing companies from training their large language models on its material and the destruction of any models trained on its work.

OpenAI and Microsoft do not dispute that its models trained on the news outlet’s content, but argue that doing so is legally permissible.

While not a party in the case, the DOJ’s statement of interest agrees, writing that “constraining LLM development under a misunderstanding of fair use doctrine would thwart such creative and scientific progress while hindering American prosperity and economic mobility.”

How does training AI work?

Copyright law can impact both the input (the training) and output (the response) of an AI model.Harvard Law professor Lawrence Lessig called copyright law “the most inefficient property system known to man,” with no good way to identify who the owner of most copyrighted work is.

If companies end up needing to clear rights before training their LLMs, Lessig said, the only companies with the resources to pay for copyright costs would be the large corporate entities like OpenAI and Microsoft.

“Training has got to be protected as fair use,” Lessig added. “It would be disastrous policy — copyright policy and innovation policy — if it were not.”

Content like articles and videos on a publication’s website is first used in the “pre-training” stage of machine learning. This is when models take in large-scale data to learn to predict which word comes next in an input, building foundational representations of language.

Then, in “post-training,” also called “reinforcement learning,” the models are fine-tuned to be less of a predictor and more of a conversational chatbot, producing human-like responses.

The New York Times suit targets both input and output, arguing that copyright restricts the ability to train on copyrighted material and requires that the end product be significantly different than the input material.

“AI should be able to train on copyrighted works, though if the output copies those works (which happens sometimes, but not often) that is more likely to be infringing,” Mark Lemley, professor of law at Stanford Law School, wrote in an email to USA TODAY.

The issue with suing for copyright infringement is that almost everything that anyone has written in the last century is under U.S. copyright, “so if models can’t train on published information or anything on the web they may not have a realistic source to train at all,” he added.

The New York Times argues that its suit is focused only on content from the paper, not all internet content.

“AI companies simply need to pay fairly for the content that makes their products possible, as copyright law requires,” New York Times spokesperson Graham James told USA TODAY.

What does present-day policy say?

In California, an Anthropic copyright case settled this summer set the precedent that it is not legal to train chatbots on pirated books because the AI company does not have legal access to those books.

In the United States at large, litigation is ongoing about copyright law. The DOJ’s statement of interest makes it clear that it has a stake in where it lands.

“This Administration will never let our Nation be at a disadvantage relative to our foreign adversaries based on a plainly incorrect understanding of copyright law,” Associate Attorney General Stanley Woodward, Jr. wrote in a post on X.

The Trump administration’s backing of technology companies sits in the shadow of a potential relationship between the government and OpenAI. CEO Sam Altman proposed handing the government a 5% stake in OpenAI this summer, raising questions about whether the government’s involvement in the case is for its own investment.

Other countries’ governments are also legislating copyright law in favor of AI. In Singapore and Japan, for instance, models have freedom to train on all copyrighted material.

Lessig predicts that if training in the United States is hindered by copyright, training could shift to those countries entirely.

“There’s a kind of race to the top or race to the bottom depending on how you see it,” he said. “Even if the United States goes against the idea of free training, that’s going to be a short-term victory because that just means people will be training elsewhere.”

Greta Reich covers the artificial intelligence industry for USA TODAY through a fellowship from the Tarbell Center for AI Journalism. Funders do not provide editorial input.

This article originally appeared on USA TODAY: Can AI train on copyrighted work? The government hopes so



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *