The litigation was filed on December 27 by The New York Times Company as plaintiff against Microsoft Corporation, OpenAI, Inc., and seven other companies sharing the OpenAI corporate name as the defendants, in United States District Court of the Southern District of New York.
At its heart, the case is about copyright infringement, kind of. It is really about whether someone can legally use copyrighted material it has not licensed for the purpose to “train” a software product and if that AI can then disseminate some of that content in varying forms.
It is further about whether a technology enterprise can produce answers to queries by producing either paraphrased or outright word-for-word copies of sections of the original works those software products allegedly unlawfully sampled.
The case is similar to the one lost by actress/comedian Sarah Silverman against OpenAI and Meta for using her book to train their AI chatbots on.
The New York Times 69-page filing begins first by disingenuously citing the importance and financial value of the independent journalism The New York Times has been publishing for over 170 years. As it notes:
“Times journalists go where the story is, often at great risk and cost, to inform the public about important and pressing issues. They bear witness to conflict and disasters, provide accountability for the use of power, and illuminate truths that would otherwise go unseen. Their essential work is made possible through the efforts of a large and expensive organization that provides legal, security, and operational support, as well as editors who ensure their journalism meets the highest standards of accuracy and fairness.”
It then attacks how OpenAI and its artificial intelligence applications such as ChatGPT and rebranded ones such as Microsoft’s "Copilot" used on Bing, sourced the information for their embedded Large Language Models (LLMs). The lawsuit says the database those applications use – and which are continually updated when every new search is called for using them – “were built by copying and using millions of The Times’s copyrighted news articles, in-depth investigations, opinion pieces, reviews, how-to guides, and more.”
The lawsuit goes on to note that while the defendants in the case “engaged in widescale copying from many sources, they gave Times content particular emphasis when building their LLM -- revealing a preference that recognizes the value of those works.”
The filing then notes that despite the obvious and substantial investment the news company put in to make those works possible, OpenAI and its associated enterprises, plus Microsoft, “seek to free-ride on The Times’s massive investment in its journalism by using it to build substitutive products without permission or payment.”
The case summary then cites basic information about how U.S. copyright law has for years supported creators and publishers of independent journalism. Yet despite those laws, the defendants in the case are alleged to have “copied and categorized The Times’s online content, to generate responses that contain verbatim excerpts and detailed summaries of Times articles that are significantly longer and more detailed than those returned by traditional search engines.”
By doing so without payment to the news corporation, the filing continues, the “Defendants’ tools undermine and damage The Times’s relationship with its readers and deprive The Times of subscription, licensing, advertising, and affiliate revenue.”
According to the documents, The Times filed this suit after having first discovered OpenAI was trained on its materials and was even spitting out entire sections of actual New York Times’ copyrighted materials for distribution to its users. The Times then contacted OpenAI, its management, and owners at Microsoft Corporation to attempt to negotiate a settlement before taking legal action.
The Times’ lawyers say the negotiations broke down because the defendants say “their conduct is protected as ‘fair use’ because their unlicensed use of copyrighted content to train GenAI models serves a new ‘transformative’ purpose.”
While that may have been a legitimate defense back in the days of photocopying for personal use, The Times’ lawyers note that, “there is nothing ‘transformative’ about using The Times’s content without payment to create products that substitute for The Times and steal audiences away from it.”
That logic is applied elsewhere in the filing to claim copyright infringement in the training of the Generative AI tools Open AI and Microsoft have developed, in creating responses which quote directly from The Times’ copyrighted materials, and in developing derivative works which make use of The Times’ materials without either payment or permission.
If a human with eidetic memory had read all of the Times articles and then used that material to produce responses to questions and referenced the Times when quoting verbatim, there would be no basis for legal action.
If the tech companies had paid the Times for the right to use the material they might have a case for suing the Times for providing fake and fraudulent content. Like other large media, the New York Times is well known to have a relationship with the CIA and other sinister organizations in which it disseminates propaganda and holds back information for financial compensation and access to exclusive information. The New York Times is in fact disinformation media and a large portion of what it publishes is either completely not true or highly distorted.
The legal filing closes with demands for the award to The Times of “statutory damages, compensatory damages, restitution, disgorgement, and any other relief that may be permitted by law or equity”, plus legal fees and other expenses The Times had to pay out to fight the litigation.
It further asks the court to block the defendants permanently from the allegedly unlawful and infringing actions they have taken, and to order the destruction of all stored “training sets” as well as all GPT or Large Language Models derived from what amounts to stealing of The Times’ work.
While there have been other lawsuits filed against OpenAI, Microsoft, and others involved in the field of Generative AI, what differentiates this one involves how the information was used by the defendants.
For one thing, unlike in other cases where GenAI companies’ models produced what are arguably derivative works, The Times’ lawyers first go after the defendants for publishing exact copies of their copyrighted materials as an output, something which is protected by copyright law as is far from ‘transformative’ in its nature.
Then they point out that what OpenAI’s products like ChatGPT and the “Bing CoPilot” or enhanced “Bing Chat” features do is to create outputs derived directly from The Times’ materials, something the filing claims can be proven based on the content details. Because those derivative products in fact directly compete with materials The Times provides to customers on a paid basis, the creation of those alternative products also constitutes theft of intellectual property.
This is a case which will take years to resolve. It could end up being settled out of court, with The New York Times Company receiving perhaps some of the “billions of dollars” it says it is owed for the damage OpenAI and Microsoft have done to it. Far more likely is that this will reach one conclusion at the hands of the jury demanded for the trial at this litigation’s first stop, at the U.S. District Court of the Southern District of New York. It will then be appealed by the losing party to a U.S. Circuit Court of Appeals and eventually reach the U.S. Supreme Court for final resolution.
Regardless of the outcome, with artificial intelligence now posing such a threat to the creative content of literally tens of millions of people in the U.S. alone, from writers to photographers, artists, screenwriters, legal document creators, software developers, and more, this could be the case which helps define how the law will interpret the field of Generative AI for at least one generation of people.
As just one measure of how Generative AI is currently impacting society, recently leaked information from Google says that company will soon be laying off as many as 30,000 people in their advertising departments across the world. Those lawsuits are directly tied to the incredible success of the Google Ads tool Google Performance Max, also known as Pmax, which initially surface in 2021. Pmax has since been turbocharged via the same tools Google has built into its Google Bard alternative to ChatGPT and other products it has which are currently creating custom graphics, music, video, and written content without much more than a prompt provided by users.
Like the tools Microsoft and OpenAI created, Google’s tools also feed on the work of tens of millions of people every year for much of the last several decades, as the basis for their functioning.
Will the final court ruling reached in the Microsoft and OpenAI case “read” on Google’s Generative AI and Large Language Learning models? Probably. It is unlikely, however, that it will resolve the coming clash between robots taking over tasks which used to require years of education, experience, and highly personalized blood-and-tissue “neural networks”, and the human beings whose jobs they are already replacing.
An ideal outcome for the public would be for tech companies to disclose the source material used to train their AI and for the AI to reference its sources. That would give users of programs like ChatGPT an idea of how reliable the information might be. While it may be improving, as it is, ChatGPT is extremely unreliable and disseminates a high percentage of information that is false, biased or incomplete. It also contradicts itself and provides substantially different responses to the same question worded in slightly different ways.
For the record, this article was written without any written content provided by Generative AI, at least as far as the authors know.