I don't like it but the Trump DoJ is right on AI and fair use

Share
I don't like it but the Trump DoJ is right on AI and fair use
Not AI this is a real photo of a tiny robot reading a book

I've written about this topic a number of times before, but in this case I'm tempted to re-use the popular meme "When the worst person in the world makes a good point," as I did awhile back when I wrote about Elon Musk's lawsuit against OpenAI for (among other things) not actually being open. In this case, the worst person is the Trump government – or, to be more specific, the Department of Justice. Obviously not everyone in the DoJ is tainted by the association with Trump, and many of them may be hard working and smart lawyers who would like nothing better than to be associated with a different administration, and I choose to believe that those are the ones who wrote the brief that the Trump government recently filed in the copyright case between the New York Times and OpenAI. In a nutshell, the brief argues that scanning news stories and other content owned by the New York Times for the purposes of training an AI model should be covered by the fair use exemption in U.S. copyright law – and, of course, implies that all content owned by corporations or other rightsholders is theoretically covered by the same exclusion. As some of you may know by now, this is a conclusion that I agree with.

To recap for those of you who aren't familiar with the case, the New York Times sued OpenAI in December of 2023, claiming that its ChatGPT model was trained on NYT content, and that the newspaper publisher should be entitled to billions of dollars in compensation as a result. The newspaper's complaint, filed in Manhattan federal court, accused OpenAI (and its then-partner Microsoft) of trying to "free-ride on The Times's massive investment in its journalism" by using it as a way to give information to users. The lawsuit added that "there is nothing 'transformative' about using The Times's content without payment to create products that substitute for The Times and steal audiences away from it." This latter statement is a crucial element in the newspaper company's argument, since one of the four factors that the courts take into account when considering a fair-use case (and the one that courts often deem to be the most important factor) is whether the infringing use could be seen as "transformative."

It's worth noting that the point about OpenAI free-riding on the Times in order to create an alternative way of giving users the same information was buttressed in the lawsuit by an exhibit the Times said showed that ChatGPT gave users the entire text of NYT articles for free when asked to do so. However, OpenAI responded that this was not the way the AI model was supposed to be used, and that the Times only managed to generate the entire text of an article by repeatedly prompting the AI with long sections of the article, and even then it took many tries before it succeeded. “Even when using such prompts, our models don’t typically behave the way The New York Times insinuates, which suggests they either instructed the model to regurgitate or cherry-picked their examples from many attempts,” OpenAI said at the time. In any case, OpenAI said that it had changed the way ChatGPT worked so that this was no longer possible. The whole point of the model, the company said, was not to reproduce the entire text of existing articles but merely to use those articles as raw material to train ChatGPT to understand text.

As it happens, I wrote about the question of AI training and fair use for the Columbia Journalism Review (where I was then the senior technology writer) two months before the Times filed its lawsuit, in a piece entitled "An AI engine scans a book. Is that copyright infringement or fair use?" At the time, publishers and content creators of all kinds were trying to figure out whether to partner with or sue OpenAI, which was the only AI model anyone was really familiar with, and the only one with a consumer portal. At the time I wrote that piece in 2023, the Times was in discussions with OpenAI about a potential licensing deal for its content, and had been for some time – and then it launched the lawsuit, something that OpenAI said it was blindsided by, since it thought negotiations were going well. In any case, in my piece for CJR I wrote:

Determining whether LLMs training themselves on copyrighted text qualifies as fair use can be difficult even for experts—not just because AI is complicated, but because the concept of fair use is, too. According to a 1990 Supreme Court ruling, the doctrine was initially intended to counterbalance decisions under the Copyright Act of 1976 that might inadvertently “stifle the very creativity which that law is designed to foster.” The US Copyright Office notes that the Act lists certain types of activity—such as criticism, comment, and news reporting—as examples that qualify under the exemption. But judges deciding such cases have to take into account four separate and in some cases competing factors: the purpose of the use and whether it is “transformative,” the nature of the copyrighted work, the amount of the work used, and what effect the use has on the market for the original.

Note: In case you are a first-time reader, or you forgot that you signed up for this newsletter, this is The Torment Nexus. Thanks for reading! You can find out more about me and this newsletter in this post. This newsletter survives solely on your contributions, so please sign up for a paying subscription or visit my Patreon, which you can find here. I also publish a daily email newsletter of odd or interesting links called When The Going Gets Weird, which is here.

Fair use is worth defending

Another real photo that is definitely not AI generated

To me, at least, fair use is a crucially important concept that often gets overlooked in discussions about copyright, and specifically about whether scanning or indexing content for search, AI use etc. qualifies as fair use. Many people seem to be under the impression that copyright law exists solely to compensate creators for the use of their content, and that only a content creator gets to say what uses they deem to be appropriate. This is simply inaccurate, and the reason many people are under this impression (I think) is that content owners and copyright holders often deliberately misrepresent the law in order to benefit from their ownership of content rights. Music publishers routinely argued in the past that downloading songs for personal use amounted to theft, despite the fact that theft is a crime and copyright infringement is not (something the US Supreme Court has spelled out multiple times) – specifically because infringement doesn't deprive the rightsholder of anything tangible in the same way that theft does.

The reason the fair use exemption was created was to counterbalance the rights of the content creator or rightsholder. Congress decided that a content creator's interest in maximizing the revenue from their creation needed to be balanced by society's interest in enhancing and promoting the exercise of free expression, education, and creativity. So the fair-use exemption is designed to allow copyrighted content to be used – by default, not just with the permission of the owner – for education, for journalism, and to create other works that might make use of some or all of the content owned by the rightsholder. As I noted in the CJR piece, a finding of fair use can be made even if the user is charging money or has a financial interest in their use of the copyrighted work, and one of the ways that can be counteracted is through the "transformative" factor. This is what happened in the Google Books case, where the company was sued for scanning millions of books, and eventually won. Here's what the judge said in that decision:

Google Books provides significant public benefits. It advances the progress of the arts and sciences, while maintaining respectful consideration for the rights of authors and other creative individuals, and without adversely impacting the rights of copyright holders. It has become an invaluable research tool that permits students, teachers, librarians, and others to more efficiently identify and locate books. It has given scholars the ability, for the first time, to conduct full-text searches of tens of millions of books. It preserves books, in particular out-of-print and old books that have been forgotten in the bowels of libraries, and it gives them new life. It generates new audiences and creates new sources of income for authors and publishers. Indeed, all society benefits.

I'm sure content owners and rightsholders would object to my comparing these two cases, in part because they would argue that there's no way that AI "creates new sources of income for authors and publishers." If anything, they would probably argue that it does the exact opposite, by "stealing" their content and then generating derivative works that fill the same need and therefore deprive them of revenue. And it may in fact do that! But again, the simple fact that authors or creators are deprived of some theoretical future revenue is not enough to quash a finding of fair use. The larger question is whether the use is transformative enough that it creates benefits for society that outweigh the potential revenue loss for creators. One potential argument might be that AI scanning of content is bad because AI ultimately deprives society of the benefits of free expression, education and creativity because no one will do those things if AI does, but that is a bit of a stretch as a legal argument, and I doubt a court would accept it.

In the DoJ filing in the NYT-OpenAI case, the brief states that the Times is seeking to narrow fair-use doctrine to exclude the training of OpenAI’s large language models, but argues that this "would be inconsistent with basic copyright law principles and severely hamper the Progress of Science and useful Arts,” as defined in the legislation. Many publications -- including the New York Times itself – are already making use of AI for journalistic purposes, and in many cases the use of AI tools could help smaller outlets compete with larger media conglomerates (like the NYT). And while it isn't relevant to a copyright case, the DoJ notes that AI models are "helping researchers across fields achieve major breakthroughs." Constraining AI development under a misunderstanding of fair use doctrine would "thwart such creative and scientific progress while hindering American prosperity and economic mobility," the brief says.

Transformation is in the name

This is also a real photo and not AI generated

I should point out that the DoJ isn't intervening in the NYT-OpenAI case solely because it has thoughts about whether indexing content for AI training is fair use – the brief also notes in the opening that the Trump administration is interested in keeping the US as a dominant player in AI, and by extension is worried that hampering some of the industry's leading entities might allow China to win (reading between the lines). But that said, I think the arguments it makes are solid, including where it says that an erroneous fair use ruling would hamper competition because "only the largest technology companies might have the capital necessary to pay licensing fees, and such licensing fees would disproportionately benefit legacy media outlets" such as the New York Times. It is not in the public’s interest, the filing adds, "for the largest technology companies to have an oligopoly on LLM training due to licensing entry barriers that function primarily as large subsidies for old mainstream media companies." Whatever the motivation, I agree.

In a recent essay on this topic, longtime tech writer Kevin Kelly – a cofounder of Wired magazine among other things – noted that during training an AI literally transforms text and images and music and other forms of human expression into a "mathematical abstraction that contains no expression." This is why large-language models are called transformers, Kelly says (in the the name ChatGPT, the GPT part stands for "generative pre-trained transformer"). Copies of books and articles and music don't exist per se inside the model, even though all of their constituent parts do in a sense – they have been converted into "tokens" or syntactical expressions that the AI recombines based on similarities, the prompt it is given, etc (basically math). In that way, Kelly says, it is a "fundamental transformation, because it goes one-way. It is an asymmetrical process: Content can be moved into latent space, but the latent space cannot be reversed to go back into the original content. It has truly been transformed."

As I wrote last year, a judge in a copyright case launched against Anthropic for using a database of pirated books for training ruled that the indexing of the books themselves could be considered fair use, but using a database of illegally acquired books was not. "Every factor but the nature of the copyrighted work favors this result," he wrote. In fact, he added that the technology at issue was "among the most transformative many of us will see in our lifetimes." The process used to convert purchased print library copies into digital library copies was justified by fair use as well, the judge ruled. Just two days later, a judge reached a similar result in Meta’s favor in the Kadrey v. Meta Platforms, Inc. case, allowing Meta’s use of works by authors, including Sarah Silverman and Ta-Nehisi Coates. Not only did the judge find that indexing the works was transformative, he also found that there was insufficient evidence of market harm.

As I noted in my previous post on this topic, there are a number of groups that believe AI indexing of content should qualify as fair use. The Library Copyright Alliance has argued that the ingestion of copyrighted works to create large language models or other AI training databases "generally is a fair use,” an opinion it also provided in a submission to the US Copyright Office's notice of inquiry on AI. Their reasoning is that if copyright maximalists were to lock down content from LLMs, it might also make that content less available to others as well. As the Alliance put it: "As champions of fair use, free speech, and freedom of information, libraries have a stake in maintaining the balance of copyright law so that it is not used to block or restrict access to information." Licensing schemes, the group wrote, "could stunt the development of this technology, and undermine its utility to researchers, students, creators, and the public."

I think it's important to distinguish in a case like this between the inputs that an AI model uses for training and the outputs that it generates. For example, some artists and copyright holders have noted that if you ask ChatGPT or any other AI to generate a work in the style of Marvel comics or Vincent Van Gogh, the AI will often create something that looks remarkably similar to an existing work, whether it's Iron Man or Starry Night. This is something the courts might take an interest in regardless of how they rule on whether the indexing of that content constitutes fair use. So the ingestion of content might be fair, but the duplication of existing works might not be.

If you liked this newsletter (even if you didn't agree with it) please consider upgrading to a paid subscription, or donating through my Patreon. Got any thoughts or comments? Feel free to either leave them here, or post them on Substack or on my website, or you can also reach me on Twitter, Threads, BlueSky or Mastodon. And thanks for being a reader.