Beta The Briev beta is out. Free on iPhone via TestFlight — install it in under a minute.

Join the beta ↗
Briev
Live
Technology

Study finds large AI diffusion models lose traceability to training data

MIT researchers discovered that as diffusion models are trained on ever larger datasets, linking their outputs to specific source material becomes increasingly impossible.

In a paper titled "Outputs of Generative Diffusion Models are Often Unattributable," MIT computer scientists Zheng Dai and David K Gifford reported that attribution of AI-generated content to individual training examples deteriorates as models scale up. Testing with ablation—removing specific images from the training corpus—showed that even after deleting works like the Mona Lisa, large diffusion models continued to recreate similar outputs.

The authors argue this "attribution decay" challenges legal efforts to trace AI creations back to copyrighted material, potentially shielding companies from liability. They note that the phenomenon also blurs the line between copying and novel creation, affecting fair-use assessments and artist compensation. Law professor James Grimmelmann highlighted that courts may need alternative methods to evaluate alleged copying in image-generation cases. The study will appear in Nature Communications on Tuesday.

Why it matters

If AI outputs can’t be linked to specific training data, copyright enforcement and artist compensation become far more difficult.

In this story

diffusion modelsattribution decaytraining datacopyright lawsuitsfair useAI-generated imagesMIT CSAILNature CommunicationsMidjourneyStable Diffusion
Get the beta ↗