1 paper
Mengyi Deng, Zhiwei Li, Xin Li +4
Although Large Language Models (LLMs) have made remarkable progress, current preference optimization methods still struggle to align directional consistency while preserving reason…