Introduction
Synthetic data—artificially generated datasets that mimic real-world information—has emerged as a powerful tool for training AI, testing systems, and preserving privacy. But when archives serve as the source material for synthetic data, questions of compliance, governance, and ethics arise. Is synthetic data a friend that protects privacy, or a foe that amplifies risks? This blog unpacks the dual role of synthetic data in the archiving ecosystem.