paper

Broken Promises of Privacy: Responding to the Surprising Failure of Anonymization

  • Authors:

📜 Abstract

Computer scientists have recently undermined our faith in the privacyprotecting power of anonymization, the name for techniques that protect the privacy of individuals in large databases by deleting information like names and social security numbers. These scientists have demonstrated that they can often “reidentify” or “deanonymize” individuals hidden in anonymized data with astonishing ease. By understanding this research, we realize we have made a mistake, labored beneath a fundamental misunderstanding, which has assured us much less privacy than we have assumed. This mistake pervades nearly every information privacy law, regulation, and debate, yet regulators and legal scholars have paid it scant attention. We must respond to the surprising failure of anonymization, and this Article provides the tools to do so.

✨ Summary

Summary

Paul Ohm argues that the prevailing belief that removing names, Social Security numbers, and other explicit identifiers can make datasets effectively anonymous is technically and legally unsound. Drawing on reidentification research, including studies involving demographic records, online search histories, and movie-rating data, the article shows that seemingly innocuous attributes can be linked with external information to identify individuals and reveal sensitive facts. The central claim is that data can often be useful or strongly anonymous, but not both at the same time.

The article contends that this problem destabilizes privacy laws and regulatory frameworks that rely on personally identifiable information (PII) as a threshold for protection or that create exemptions for “anonymized” data. Ohm recommends moving away from a binary distinction between personal and nonpersonal data, abandoning the assumption that anonymization is a reliable privacy guarantee, and using terminology that describes privacy-protective effort without implying successful anonymity. He proposes assessing reidentification risk contextually, taking account of the data-handling technique, whether disclosure is public or restricted, the amount and sensitivity of the data, the availability of outside information, and the likely incentives and capabilities of adversaries.

The article rejects several incomplete responses, including relying solely on after-the-fact compensation, waiting for technical solutions to eliminate the problem, or prohibiting reidentification. It instead supports a combination of comprehensive privacy regulation, sector-specific rules, restrictions on data retention and dissemination, controlled access, auditing, aggregation, and privacy-preserving techniques such as differential privacy. The conclusion emphasizes that regulation must balance the risks of information flows against competing interests such as research, innovation, security, and free expression. (studylib.net)

Influence

Subsequent scholarship has treated the article as a major legal articulation of the limits of deidentification. Later work on “functional anonymisation” identifies Ohm’s argument as perhaps the most influential anti-anonymization position and develops a contextual alternative that evaluates anonymization within a broader data environment. (sciencedirect.com) Princeton researchers also state that the article helped spur legal and policy debate over how to respond to reidentification research, including consideration of differential privacy and contractual or legal controls on data use. (blog.citp.princeton.edu) The paper continues to be cited in discussions of health-data privacy, big-data governance, deidentification, differential privacy, and the legal meaning of personal data. (georgetownlawtechreview.org)