An Ethical Examination of the Internet Census 2012 Dataset: A Menlo Report Case Study
📜 Abstract
In 2012, an anonymous individual used basic techniques such as default or no passwords on Internet access equipment to get unauthorized access and install a scanning botnet on hundreds of thousands of commodity network devices around the world. Making efforts to minimize potential harm, this individual still took actions that would never be sanctioned by an academic ethics review board (and which violated computer trespass statutes around the globe), to perform an unprecedented internet-wide scan at a scale and speed never seen before. Then, some results and the raw data were published for anyone to access, raising a host of ethical and legal questions. One central question is whether researchers who, but for the illegally obtained data, could not ethically or legally perform the same experiment or produce the same data themselves should use that data. Or might the lack of community condemnation for performing such potentially illegal and unethical experiments create a situation where researchers are effectively encouraging law breaking by those willing to risk getting caught to create data that is otherwise not justifiable to create, simply to allow researchers to get around ethical restrictions? In this paper we examine this event in the context of the guiding principles outlined in the Menlo Report in an attempt to better understand the ethical implications of such actions.
✨ Summary
The paper applies the Menlo Report’s ethical framework to the Carna botnet and the Internet Census 2012 dataset. It identifies device owners, network operators, the anonymous botnet creator, criminals exploiting the same vulnerabilities, and subsequent researchers as relevant stakeholders. The analysis concludes that the operation violated respect for persons through lack of consent and violated respect for law through unauthorized access, while the operator’s efforts to limit technical damage did not resolve the underlying ethical problems. It also examines whether researchers should reuse data obtained through potentially illegal and unethical conduct, emphasizing the tension between unique research value, reproducibility, public interest, privacy, and the risk of incentivizing future unlawful data collection.
Documented follow-on work shows both conceptual and technical uptake. Thomas et al.’s 2017 study of research using datasets of illicit origin cites this paper as a case study and places the Internet Census within a broader analysis of ethical safeguards, harms, benefits, and research oversight. (cl.cam.ac.uk) Separately, Nikolić et al.’s 2014 Hadoop and Pig paper used the Internet Census 2012 dataset as the basis for a distributed large-scale data-analysis platform, demonstrating concrete academic reuse of the dataset discussed in this paper. (researchgate.net) The sources reviewed document academic follow-on research, but do not establish a specific industry deployment or policy adoption attributable directly to this paper.