3 papers
cs.LG2026
Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety
Domenic Rosati, Ali Dadsetan, Hong Huang +5
A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing su…
cs.LG2025
SoftAdaClip: A Smooth Clipping Strategy for Fair and Private Model Training
Dorsa Soleymani, Ali Dadsetan, Frank Rudzicz
Differential privacy (DP) provides strong protection for sensitive data, but often reduces model performance and fairness, especially for underrepresented groups. One major reason…
cs.LG2024
Can large language models be privacy preserving and fair medical coders?
Ali Dadsetan, Dorsa Soleymani, Xijie Zeng +1
Protecting patient data privacy is a critical concern when deploying machine learning algorithms in healthcare. Differential privacy (DP) is a common method for preserving privacy…