2 papers
cs.LG2026
Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
Ren-Wei Liang, Chin-Ting Hsu, Chan-Hung Yu +6
Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessive refusals, while permissive mode…
cs.LG2024
Adversarial Robustness Overestimation and Instability in TRADES
Jonathan Weiping Li, Ren-Wei Liang, Cheng-Han Yeh +4
This paper examines the phenomenon of probabilistic robustness overestimation in TRADES, a prominent adversarial training method. Our study reveals that TRADES sometimes yields dis…