Tuesday, 25 August 2026 SourcesAbout🌓
🇬🇧 UK ▾
BREAKING
Technology

'Unlicensed, unrestricted AI training could destroy the ecosystem for books' — quote of the day by the Authors Guild on the sourcing of training data

TechRadar ·
'Unlicensed, unrestricted AI training could destroy the ecosystem for books' — quote of the day by the Authors Guild on the sourcing of training data

To achieve any level of competency, large language models (LLMs) need ample data for sufficient training.

AI companies have looked to various sources to mine this information, including content publicly available on the internet, synthetic data generated from other AI models, and printed literature.

Reading difficulties Prompted by news that AI companies were allegedly using books from pirate ebook sites to build their LLMs, writers, authors, and publishers publicly called out this deeply worrying process.

Quote of the day This article is part of TechRadar Pro's QOTD project to provide an insight into the minds of the brightest and most recognized figures in the technology industry today and in years gone by.

Read the full series here .

The professional organization known as the Authors Guild responded to various stories about AI companies scanning books to train their AI models (both illegally and legally) with incredibly comprehensive guidelines on AI licensing .

This document covered the various manifestations of the use of published works by AI companies, including its legal perspective on the legitimacy of using such works.

It also highlighted that the continued data harvesting processes would risk destroying the ecosystem for books that currently exists.

Book buying The use of books by AI companies is an ongoing concern.

But in recent months the focus has pivoted to those that buy, scan – and destroy – books on an industrial scale.

For example, court documents revealed the existence of ' Project Panama ' inside Anthropic.

This is a scheme in which the company aims to "destructively scan all the books in the world" and used a codename because "we don’t want it to be known that we are working on this.".

To train Claude, Anthropic had to procure a large and high-quality dataset, so it set out to purchase books on an industrial scale because of the relatively high-quality nature of the writing compared with, say, writing found online.

Read the full article on TechRadar ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.techradar.com — the content belongs to TechRadar.

More from TechRadar

See all ›

More in Technology

See all ›