A new preprint argues that current methods for assessing the anonymity of synthetic data fail to account for model-centric privacy attacks, potentially exposing sensitive personal information.

The study, published on arXiv S1, contends that releasing trained models or synthetic datasets can still pose privacy risks, as these methods are often evaluated at the dataset level rather than considering the underlying generative model.

According to the preprint, the General Data Protection Regulation (GDPR) definitions of personal data and anonymization need to be reinterpreted under the assumption that trained models are accessible for interaction or querying. The authors map identifiability risks to privacy attacks across various threat settings and argue that synthetic data techniques alone do not ensure sufficient anonymization.

The preprint compares two commonly used mechanisms with synthetic data: Differential Privacy (DP) and Similarity-based Privacy Metrics (SBPMs). The authors conclude that while DP can offer robust protections against identifiability risks, SBPMs lack adequate safeguards.

The study was presented at the 25th Workshop on Privacy in the Electronic Society, WPES 2026, part of ACM CCS 2026.

The findings come amid heightened scrutiny of AI risks, with companies like Apple tightening permissions for autonomous AI agents S3.

The preprint has not been peer-reviewed, and the authors may have commercial interests.