2 papers
cs.LG2025
LLM Safety Alignment is Divergence Estimation in Disguise
Rajdeep Haldar, Ziyi Wang, Qifan Song +2
We present a theoretical framework showing that popular LLM alignment methods, including RLHF and its variants, can be understood as divergence estimators between aligned (safe or…
stat.ML2024
Adversarial Vulnerability as a Consequence of On-Manifold Inseparibility
Rajdeep Haldar, Yue Xing, Qifan Song +1
Recent works have shown theoretically and empirically that redundant data dimensions are a source of adversarial vulnerability. However, the inverse doesn't seem to hold in practic…