Search papers, labs, and topics across Lattice.
This position paper critiques the evolving standards of privacy in the context of synthetic data usage in machine learning, highlighting a shift from a focus on residual inference risk to an emphasis on appearance-based properties. The authors conducted an empirical analysis of recent publications, revealing that many studies utilize synthetic data in privacy-sensitive applications without clearly defined threat models or privacy claims. They advocate for a redefinition of privacy as an explicit, evidence-based claim, urging ML venues to enforce norms that ensure privacy assertions are clearly articulated and verifiable.
Privacy in synthetic data is often treated as an implicit assumption rather than an explicit, testable claim, leading to uneven protections for sensitive information.
Synthetic data has become a common component of machine learning research. While widely adopted, its use in privacy-sensitive contexts has quietly shifted from a claim of residual inference risk under stated assumptions to an appearance-based property inferred from data generation itself. In this position paper, we argue that this shift reflects an implicit change in community standards for what counts as sufficient privacy evidence, rather than a misunderstanding of well-established privacy principles. Drawing on an empirical analysis of recent publications across major ML venues, we show that synthetic data is frequently used in privacy-sensitive settings without explicit articulation of threat models, inference risks, or falsifiable privacy claims. As a result, privacy assurance often remains implicit, difficult to verify, and unevenly distributed, with heightened exposure for rare and minority records. We argue for treating privacy as an explicit, evidence-based scientific claim and recommend that ML venues adopt norms requiring privacy-relevant assertions to be clearly scoped, testable, and contestable.