Opinion | Even Millions Of Stolen Books Cannot Satisfy Ravenous A.I. Chatbots – The New York Times
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

PRIME

Get ready for Prime Big Deal Days — try Prime free

Exclusive member deals on October 6–7, plus fast free delivery. Cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

An opinion piece in the New York Times argues that even a large collection of stolen books would be insufficient for training AI chatbots. The claim’s evidence and targets are unconfirmed, and the legal basis is unclear.

The New York Times has published an opinion article asserting that even millions of stolen books would be insufficient to satisfy the data demands of AI chatbots as detailed in the original analysis. The piece highlights ongoing disputes over copyright and data sourcing but provides no concrete evidence, legal findings, or specific AI systems involved. This claim intensifies debates about the provenance of training data and the scale of resources needed to develop advanced language models, which are discussed in detail in the original analysis.

The opinion piece, published in the New York Times, suggests that the volume of books—described as millions—used or potentially used in AI training is insufficient to meet the ravenous data requirements of current language models. The headline implies that these books are stolen, raising questions about copyright infringement and unauthorized data collection. However, the article offers no specific evidence or documentation to support the claim, such as court rulings, dataset disclosures, or company responses.

It remains unclear which AI developers or models are referenced, or whether the books were obtained illegally or through licensed sources. The phrase “stolen” appears to be an opinion rather than a verified legal judgment, as explored in the original analysis. The article emphasizes the dispute over data rights and the scale of data needed for effective AI training, but the specifics are not provided.

At a glance
analysisWhen: published August 2026
The developmentA New York Times opinion article asserts that millions of stolen books cannot satisfy AI chatbots’ data requirements, sparking debate over data sourcing and copyright issues.

Implications for AI Data Sourcing and Copyright

This discussion matters because it underscores the conflicting interests in AI development: access to large datasets, copyright protections, and the ethical considerations of data collection. If AI companies rely on unauthorized or stolen materials, they risk legal repercussions and public backlash. Conversely, the claim that even extensive collections of such material are insufficient raises questions about the efficacy of current training methods and the future demand for data. For authors, publishers, and consumers, these issues influence copyright enforcement, access to information, and model transparency.

Amazon

AI training dataset books

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Ongoing Debates Over Data and Copyright in AI

The use of large-scale text datasets in AI training has long been controversial, with debates centering on whether copyrighted works are used without permission. Past incidents have involved allegations of unauthorized data scraping and the use of publicly available versus licensed materials. The scale of data needed for training models like GPT-4 and similar systems has grown significantly, fueling concerns over copyright infringement and data provenance.

The headline’s reference to “millions of stolen books” echoes ongoing disputes but does not specify which datasets or companies are involved. Previous disclosures from AI firms have shown reliance on a mixture of licensed, public domain, and scraped data, but detailed legal or dataset documentation remains scarce.

“The claim that millions of stolen books are used in training is an unverified assertion that highlights ongoing legal and ethical concerns.”

— Thorsten Meyer, AI researcher

Amazon

AI language model training books

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Claims and Lack of Evidence

It is not yet clear which specific books, datasets, or AI systems the opinion refers to. The headline provides no details on the source of the “millions” of books, whether they were obtained legally or illegally, or if any court has ruled on the matter. The claim that these books are “stolen” remains an opinion-based assertion without supporting documentation or legal findings. The extent to which these materials are used in actual AI training is also unconfirmed.

Amazon

large data storage for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Awaiting Full Article and Data Disclosures

Further clarification will depend on access to the full opinion column, any cited lawsuits, dataset disclosures, or company statements. Monitoring responses from AI developers, rights holders, and legal authorities will be crucial. Future developments may include legal rulings, dataset transparency initiatives, or policy debates on data rights and AI training practices.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Does the headline prove that millions of books were stolen for AI training?

No. The headline is an opinion assertion without supporting evidence or legal rulings. It suggests a possibility but does not confirm actual theft or illegal use.

Which AI companies are involved in this claim?

The headline does not specify any companies or AI models. It remains a general statement without attribution to particular developers or systems.

Why would AI developers use books in training?

Books can provide long-form language, structured arguments, and diverse subject matter, which are valuable for training generative language models. However, the source and legality of such data are often contested.

The primary concern is whether the data was obtained legally through licensing or fair use, or illegally via unauthorized scraping or copying. Courts have yet to definitively rule on many such cases, and transparency about datasets remains limited.

Source: ThorstenMeyerAI.com

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Artificial Intelligence Combines Deep Strikes And Jamming Into One System

New AI-enabled military systems now combine deep strike and electronic jamming, transforming modern warfare strategies and penetration tactics.

GLM-5.3-Flash: A Low-Cost AI Agent Engine With Potential And Problems

Z.ai releases GLM-5.3-Flash, a 320-billion-parameter multimodal model with open weights, designed for efficient agent workflows but with hosting limitations.

The Nordics: Protect the Worker, Not the Job

Exploring how Nordic countries prioritize worker security over job preservation, enabling smoother transitions amid automation and economic change.

AI Compression And Quantization: Powering The Next-Gen Local LLMs

New advancements in AI quantization, including trained-in quantization and dynamic mixed-precision, enable efficient local large language models in 2026.