A three-year-old copyright lawsuit between The New York Times and OpenAI has new unredacted details.
Top Microsoft executives privately described AI training as theft. OpenAI leadership said its models threaten journalists' work.
The companies allegedly bypassed paywalls to collect content. They stripped copyright notices from their training data.
Microsoft's Copilot caused click-through rates for The New York Times to drop by 93%. Internal documents call this a doom loop that hurts both the company and the web.
OpenAI datasets contain over 91,692 copies of NYT works. A Common Crawl dataset included more than 2 million documents from nytimes.com alone.
Researchers allegedly planned hacks to get around paywalls without detection. They built datasets like WebText that relied heavily on scraped news content.
Why it matters
Publishers lose money when companies use their work without permission or payment.



