Why AI Models Cannot Simply Delete Your Personal Data on Request
When personal data is used to train an AI model, it does not remain as a discrete, removable record but instead becomes distributed across billions of numerical parameters throughout the model's weights. Privacy laws like the GDPR's right to erasure were designed around traditional databases where data can be located and deleted, an assumption that does not hold for trained neural networks. The only guaranteed method — retraining the model from scratch without the requested data — is prohibitively expensive, costing potentially millions of dollars per run and taking weeks to complete. Researchers are developing 'machine unlearning' techniques, such as gradient ascent and influence functions, that attempt to make a model behave as if it never saw specific data without full retraining. However, these methods are approximations and cannot offer the same provable guarantees as complete retraining, leaving a significant gap between legal data deletion rights and current technical reality.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in