15 papers
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance
Yuchun Fan, Bei Li, Peiguang Li +9
Reinforcement learning has proven effective for enhancing multi-step reasoning in large language models (LLMs), yet its benefits have not fully translated to multilingual contexts.…
DaPT: A Dual-Path Framework for Multilingual Multi-hop Question Answering
Yilin Wang, Yuchun Fan, Jiaoyang Li +5
Retrieval-augmented generation (RAG) systems have made significant progress in solving complex multi-hop question answering (QA) tasks in the English scenario. However, RAG systems…
Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
Yongyu Mu, Jiali Zeng, Fandong Meng +2
Through encouraging self-exploration, reinforcement learning from verifiable rewards (RLVR) has significantly advanced the mathematical reasoning capabilities of large language mod…
GRAM: A Generative Foundation Reward Model for Reward Generalization
Chenglong Wang, Yang Gan, Yifu Huo +8
In aligning large language models (LLMs), reward models have played an important role, but are standardly trained as discriminative models and rely only on labeled human preference…
Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
Chenglong Wang, Yifu Huo, Yang Gan +10
Previous methods evaluate reward models by testing them on a fixed pairwise ranking test set, but they typically do not provide performance information on each preference dimension…
GRAM-R: Self-Training Generative Foundation Reward Models for Reward Reasoning
Chenglong Wang, Yongyu Mu, Hang Zhou +10
Significant progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs towards generalist reward models. Despite this trend, devel…