Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization
Zhihao Liu, Yifan Wu, Jian Lou +3
Safety alignment for large language models (LLMs) aims to reduce harmful or unsafe behavior while preserving general utility. However, recent findings reveal that alignment effects…
cs.AI2026
Confidence-Aware Alignment Makes Reasoning LLMs More Reliable
Kejia Chen, Jiawen Zhang, Yihong Wu +5
Large reasoning models often reach correct answers through flawed intermediate steps, creating a gap between final accuracy and reasoning reliability. Existing alignment strategies…