3 papers
cs.CL2025
AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
Ruibo Deng, Duanyu Feng, Wenqiang Lei
Offline preference optimization offers a simpler and more stable alternative to RLHF for aligning language models. However, their effectiveness is critically dependent on ranking a…
cs.AI2025
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
Youcheng Huang, Bowen Qin, Chen Huang +3
Large Reasoning Models (LRMs) have demonstrated remarkable problem-solving abilities in mathematics, as evaluated by existing benchmarks exclusively on well-defined problems. Howev…
cs.CL2025
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
Youcheng Huang, Chen Huang, Duanyu Feng +2
Understanding the inner workings of Large Language Models (LLMs) is a critical research frontier. Prior research has shown that a single LLM's concept representations can be captur…