1 citations · 2 across the 4 of their papers we have counts for
1 paper · 1 filter
Yuhan Huang, Huanran Chen, Yinpeng Dong
Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alignment is often fragile under s…