4 papers
AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
Ruibo Deng, Duanyu Feng, Wenqiang Lei
Offline preference optimization offers a simpler and more stable alternative to RLHF for aligning language models. However, their effectiveness is critically dependent on ranking a…
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
Youcheng Huang, Bowen Qin, Chen Huang +3
Large Reasoning Models (LRMs) have demonstrated remarkable problem-solving abilities in mathematics, as evaluated by existing benchmarks exclusively on well-defined problems. Howev…
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
Youcheng Huang, Chen Huang, Duanyu Feng +2
Understanding the inner workings of Large Language Models (LLMs) is a critical research frontier. Prior research has shown that a single LLM's concept representations can be captur…
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
Duanyu Feng, Bowen Qin, Chen Huang +3
The success of the reward model in distinguishing between responses with subtle safety differences depends critically on the high-quality preference dataset, which should capture t…