2 papers
cs.LG2026
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
Eric Hanchen Jiang, Weixuan Ou, Run Liu +8
Safety alignment of large language models currently faces a central challenge: existing alignment techniques often prioritize mitigating responses to harmful prompts at the expense…
cs.LG2026
SERL: Self-Examining Reinforcement Learning on Open-Domain
Weixuan Ou, Yanzhao Zheng, Shuoshuo Sun +7
Reinforcement Learning (RL) has been shown to improve the capabilities of large language models (LLMs). However, applying RL to open-domain tasks faces two key challenges: (1) the…