Newly filed court documents in The New York Times’ copyright lawsuit against OpenAI and Microsoft allege that OpenAI president Greg Brockman was told about a method for bypassing the newspaper’s paywall and responded approvingly, while a Microsoft executive privately described the company’s data practices as “an astonishing theft of unprecedented proportions.”
The Financial Times reports that the Times and other publishers are seeking billions of dollars in damages, arguing OpenAI and Microsoft infringed copyright by training AI models on their journalism without permission. Lawyers for the Times filed the new evidence Thursday in the case, which centers on whether AI systems built on copyrighted articles constitute theft or fall under fair use protections.
The filing states that “defendants repeatedly copied millions of . . . copyrighted articles in their entirety without permission to produce substitutive commercial AI products.”
The Paywall Exchange
According to the filing, an OpenAI employee told Brockman about “a hack to get around NY Times paywall.” Brockman’s alleged response: “ah nice.”
The filing also references writing attributed to Brockman from around 2017, in which he described being “deeply motivated by the gazillions” while considering how to commercialize OpenAI’s technology.
OpenAI maintains that training its models on publicly available content is protected fair use, arguing its systems transform rather than reproduce original material. But the filing quotes OpenAI’s head of ChatGPT describing AI products as an “existential threat” to publishers, saying such products “are largely substitutive, period [and] will get more and more substitutive as they get better.”
Microsoft’s Internal Warnings
Brent Hecht, Microsoft’s director of applied science, is quoted in the filing calling the training of large language models on copyrighted content “an astonishing theft of unprecedented proportions.” Hecht is also quoted warning that a win for OpenAI and Microsoft would “make a complete mockery of the idea of ‘fair use.'”
Microsoft responded that “these comments reflect one employee’s individual perspective, are not a legal analysis.”
Separately, Microsoft’s own data cited in the filing shows click-through rates to underlying content were 83% to 93% lower for users of its AI answer tools compared with users of a traditional search engine.
Nadella’s Testimony
Microsoft chief executive Satya Nadella testified in the case that he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall.” He also testified about an agreement allowing Microsoft to require OpenAI to retrain its models under certain conditions, saying he would have required OpenAI “to retrain its models” over the paywall issue.
Microsoft said Nadella’s testimony “spoke to broad principles and changes under way in how people find and consume information” and was not a “conclusion about copyright questions.”
Evan Swarztrauber, principal of technology policy firm CorePoint Strategies, said Microsoft and OpenAI’s public legal defense is that scraping copyrighted material for training data is fair use and not a competitive threat to content creators, while their executives privately described the opposite. “In private, their executives admitted the exact opposite: that their conduct is theft that threatens the livelihoods of publishers and other copyright holders,” Swarztrauber said.
You must be logged in to post a comment Login