synthetic-data-safe-replacement-for-real-data

OUT derived (depth 1)

Created 2026-06-21T10:01:28+00:00

Synthetic data from generative models can safely replace real training data at scale, enabling privacy-preserving ML pipelines and unlimited data augmentation without degradation.

Justifications

SL — Privacy-preserving synthetic generation and decentralized training suggest synthetic data is viable, but model collapse from uncurated synthetic data undermines this — currently OUT

Antecedents (all must be IN):

  • IN gan-synthetic-medical-imaging-privacy — GANs generate synthetic medical images (MRI, PET) to overcome patient privacy barriers that limit access to real medical imaging data
  • IN ml-federated-learning-decentralized — Federated learning decentralizes training across user devices, preserving privacy by not sending raw data to a central server (e.g., Google Gboard)

Unless (any of these IN defeats this justification):

  • IN ml-model-collapse-synthetic-data — Model collapse is the degradation that occurs when models train on uncurated synthetic data or outputs of prior model versions, also called 'model autophagy disorder (MAD)'