Publications (15)
Control, Confidentiality, and the Right to be Forgotten
Aloni Cohen, Adam Smith, Marika Swanberg +1
Recent digital rights frameworks give users the right to delete their data from systems that store and process their personal information (e.g., the "right to be forgotten" in the…
From Soft Classifiers to Hard Decisions: How fair can we be?
Ran Canetti, Aloni Cohen, Nishanth Dikkala +3
A popular methodology for building binary decision-making classifiers in the presence of imperfect information is to first construct a non-binary "scoring" classifier that is calib…
Barriers to Counterfactual Credit Attribution for Autoregressive Models
Aloni Cohen, Chenhao Zhang
Generative AI disrupts the practice of giving credit to work that came before. Ideally, a generative model would give credit to any work on which its output depends in a significan…
Attacks on Deidentification's Defenses
Aloni Cohen
Quasi-identifier-based deidentification techniques (QI-deidentification) are widely used in practice, including -anonymity, -diversity, and -closeness. We present three…
Watermarking Language Models for Many Adaptive Users
Aloni Cohen, Alexander Hoover, Gabe Schoenbach
We study watermarking schemes for language models with provable guarantees. As we show, prior works offer no robustness guarantees against adaptive prompting: when a user queries a…
Census TopDown: The Impacts of Differential Privacy on Redistricting
Aloni Cohen, Moon Duchin, JN Matthews +1
The 2020 Decennial Census will be released with a new disclosure avoidance system in place, putting differential privacy in the spotlight for a wide range of data users. We conside…
Linear Program Reconstruction in Practice
Aloni Cohen, Kobbi Nissim
We briefly report on a successful linear program reconstruction attack performed on a production statistical queries system and using a real dataset. The attack was deployed in tes…
Properties of Effective Information Anonymity Regulations
Aloni Cohen, Micah Altman, Francesca Falzon +2
A firm seeks to analyze a dataset and to release the results. The dataset contains information about individual people, and the firm is subject to some regulation that forbids the…
A Machine Learning Theory Perspective on Strategic Litigation
Melissa Dutz, Han Shao, Avrim Blum +1
Strategic litigation involves bringing a case to court with the goal of having an impact beyond resolving the particular dispute at hand. In a common law system, one way a case may…
Can the Government Compel Decryption? Don't Trust -- Verify
Aloni Cohen, Sarah Scheffler, Mayank Varia
If a court knows that a respondent knows the password to a device, can the court compel the respondent to enter that password into the device? In this work, we propose a new approa…
Towards Formalizing the GDPR's Notion of Singling Out
Aloni Cohen, Kobbi Nissim
There is a significant conceptual gap between legal and mathematical thinking around data privacy. The effect is uncertainty as to which technical offerings adequately match expect…
Private PAC Learning May be Harder than Online Learning
Mark Bun, Aloni Cohen, Rathin Desai
We continue the study of the computational complexity of differentially private PAC learning and how it is situated within the foundations of machine learning. A recent line of wor…
Protecting the Undeleted in Machine Unlearning
Aloni Cohen, Refael Kohen, Kobbi Nissim +1
Machine unlearning aims to remove specific data points from a trained model, often striving to emulate "perfect retraining", i.e., producing the model that would have been obtained…
Understanding and Mitigating the Impacts of Differentially Private Census Data on State Level Redistricting
Christian Cianfarani, Aloni Cohen
Data from the Decennial Census is published only after applying a disclosure avoidance system (DAS). Data users were shaken by the adoption of differential privacy in the 2020 DAS,…
Blameless Users in a Clean Room: Defining Copyright Protection for Generative Models
Aloni Cohen
Are there any conditions under which a generative model's outputs are guaranteed not to infringe the copyrights of its training data? This is the question of "provable copyright pr…