1 citations · 1 across the 16 of their papers we have counts for
1 paper · 2 filters
Yixu Wang, Yang Yao, Xin Wang +4
Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a standard prompt but comply when t…