Extracting Training Data from Diffusion Models

Image diffusion models such as DALL-E 2, Imagen, and Stable Diffusion have attracted significant attention due to their ability to generate high-quality synthetic images.

In this work, authors show that diffusion models memorize individual images from their training data and emit them at generation time. With a generate-and-filter pipeline, we extract over a thousand training examples from stateof-the-art models, ranging from photographs of individual people to trademarked company logos. They also train hundreds of diffusion models in various settings to analyze how different modeling and data decisions affect privacy. Overall, the results show that diffusion models are much less private than prior generative models such as GANs, and that mitigating these vulnerabilities may require new advances in privacy-preserving training.

Download

February 21, 2023

Publication

Global

Arxiv

AI / research

Extracting Training Data from Diffusion Models

Support GDPR buzz!

Get notified about news & resources

Extracting Training Data from Diffusion Models

Related resources

EDPB Study on Secondary Use of Personal Data for Scientific Research

Algorithmic Risks Report

The Playbook: Data Sharing for Research

Support GDPR buzz!

Get notified about news & resources