1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Mengyi Deng, Zhiwei Li, Xin Li +4
Although Large Language Models (LLMs) have made remarkable progress, current preference optimization methods still struggle to align directional consistency while preserving reason…