
The development of generative AI relies on large-scale corpora that often contain personal information, making anonymization central to data-protection compliance. In comparative law, anonymization generally determines whether personal-information protection rules apply, with the key inquiry turning on identifiability and the reasonable likelihood of re-identification. In generative AI, however, parameterized learning, dynamic content generation, and long-term interaction have transformed anonymization from a static processing outcome into a probabilistic and context-dependent risk condition. This creates three main challenges: a mismatch between formal anonymization and substantive risk, which may lead to identity disclosure or sensitive-attribute inference; the erosion of anonymization’s function as a regulatory boundary, which may enable regulatory circumvention and undermine public trust; and uncertainty over responsibility allocation, which may hinder risk prevention, enforcement, and remedies. Existing rules offer only limited responses to residual risks, jointly produced risks, and dynamic changes in model operation. Accordingly, anonymization should be reconceptualized as an ongoing governance framework based on context-sensitive identifiability standards, tiered legal consequences, responsibility allocation, and lifecycle oversight, thereby reconciling personal-information protection with AI development.
generative artificial intelligence; anonymization; risk; re-identification risk; risk-based governance; legal determination