activity
20202026
most citedAlgorithms that Approximate Data Removal: New Results and Limitations

3 citations · 5 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Evaluation without Generation: Non-Generative Assessment of Harmful Model Specialization with Applications to CSAM

Vinith M. Suriyakumar, Ayush Sekhari, Lena Stempfle +5

Auditing the fine-tunes of open-weight generative models for harmful specialization has become a new governance challenge for model hosting platforms. The standard toolkit, generat…

cs.LG2025

When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment

Yuxin Xiao, Sana Tonekaboni, Walter Gerych +2

Large language models (LLMs) can be prompted with specific styles (e.g., formatting responses as lists), including in malicious queries. Prior jailbreak research mainly augments th…

cs.LG2025

Layered Unlearning for Adversarial Relearning

Timothy Qian, Vinith Suriyakumar, Ashia Wilson +1

Our goal is to understand how post-training methods, such as fine-tuning, alignment, and unlearning, modify language model behavior and representations. We are particularly interes…

cs.LG2024

Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models

Vinith M. Suriyakumar, Rohan Alur, Ayush Sekhari +2

Text-to-image diffusion models rely on massive, web-scale datasets. Training them from scratch is computationally expensive, and as a result, developers often prefer to make increm…

cs.LG2022

Private Multi-Winner Voting for Machine Learning

Adam Dziedzic, Christopher A Choquette-Choo, Natalie Dullerud +6

Private multi-winner voting is the task of revealing -hot binary vectors satisfying a bounded differential privacy (DP) guarantee. This task has been understudied in machine lea…

cs.LG2020

Chasing Your Long Tails: Differentially Private Prediction in Health Care Settings

Vinith M. Suriyakumar, Nicolas Papernot, Anna Goldenberg +1

Machine learning models in health care are often deployed in settings where it is important to protect patient privacy. In such settings, methods for differentially private (DP) le…