collaborators

8 papers

cs.LG2025

MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking

Yizhou Zhao, Zhiwei Steven Wu, Adam Block

Watermarking aims to embed hidden signals in generated text that can be reliably detected when given access to a secret key. Open-weight language models pose acute challenges for s…

cs.HC2025

Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks

Luke Guerdan, Devansh Saxena, Stevie Chancellor +2

Data scientists often formulate predictive modeling tasks involving fuzzy, hard-to-define concepts, such as the "authenticity" of student writing or the "healthcare need" of a pati…

cs.LG2025

Enhancing One-run Privacy Auditing with Quantile Regression-Based Membership Inference

Terrance Liu, Matteo Boglioni, Yiwei Fu +3

Differential privacy (DP) auditing aims to provide empirical lower bounds on the privacy guarantees of DP mechanisms like DP-SGD. While some existing techniques require many traini…

stat.ML2025

Generate-then-Verify: Reconstructing Data from Limited Published Statistics

Terrance Liu, Eileen Xiao, Adam Smith +2

We study the problem of reconstructing tabular data from aggregate statistics, in which the attacker aims to identify interesting claims about the sensitive data that can be verifi…

cs.LG2025

Winning the MIDST Challenge: New Membership Inference Attacks on Diffusion Models for Tabular Data Synthesis

Xiaoyu Wu, Yifei Pang, Terrance Liu +1

Tabular data synthesis using diffusion models has gained significant attention for its potential to balance data utility and privacy. However, existing privacy evaluations often re…

cs.LG2025

Improving the Convergence of Private Shuffled Gradient Methods with Public Data

Shuli Jiang, Pranay Sharma, Zhiwei Steven Wu +1

We consider the problem of differentially private (DP) convex empirical risk minimization (ERM). While the standard DP-SGD algorithm is theoretically well-established, practical im…