Tuesday, 18 August 2026 SourcesAbout🌓
🇬🇧 UK ▾
BREAKING
Technology

AI models get convenient amnesia about source material as they grow, MIT boffins find

The Register ·
AI models get convenient amnesia about source material as they grow, MIT boffins find

The process of training an AI model becomes a paradox at scale – the more it remembers, the less it remembers about the source of its memories.

MIT computer scientists went looking for a way to attribute AI model output to specific training data, in the hope that understanding could inform AI regulation.

What they found, described in a paper titled, "Outputs of Generative Diffusion Models are Often Unattributable," looks like it will actually make regulation more difficult.

Scientific journal Nature Communications will publish the paper on Tuesday.

The authors, Zheng Dai and David K Gifford, affiliated with MIT's Computer Science & Artificial Intelligence Laboratory (CSAIL), note that diffusion models like Midjourney and Stable Diffusion have become widely used tools for generating artifacts including images, videos, and audio.

Diffusion models have also attracted lawsuits from artists who argue that copies of their work included in training data have enabled AI models to reproduce their output and artistic style.

In one such ongoing copyright case from 2023, Andersen et al. v.

Stability AI Ltd, the plaintiffs have been trying to convince the court to make defendant Midjourney provide the datasets used to train its models.

The plaintiffs allege that Midjourney made its training datasets "by scraping images associated with specific artists’ names for the express purpose of enabling its model to mimic those artists’ expressive content." Being able to attribute model output to the content they ingested during training would help people understand how models function and would have various applications "including machine unlearning, data poisoning, model interpretability, fairness, and privacy," Dai and Gifford wrote in their paper.

"Furthermore, given the contemporary adoption of these models for creative and commercial purposes, attributability also carries ethical, policy, financial, and legal implications." But as it turns out, attributing model output to a specific input becomes more difficult as models get larger.

"Here we show that attribution, characterized as the task of locating a part of the training data that can be held responsible for a generated sample, can become impossible if a model is trained on a sufficiently large corpus of data," the authors state.

"We find that the more data a model is trained on, the less attributable its generated samples become, a phenomenon we henceforth refer to as attribution decay." The authors tested this by removing specific training data through a process called ablation.

And the result of this testing showed that for very large models, you could take away, for example, the image of the Mona Lisa or all of Leonardo Da Vinci's work — yet the model could still reproduce that image or style.

Read the full article on The Register ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.

More from The Register

See all ›

More in Technology

See all ›