5 papers
CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
Guofu Xie, Yunsheng Shi, Hongtao Tian +2
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning abilities of Large Language Models (LLMs) by using rule-based binary feedback. However, current RLV…
Merge and Guide: Unifying Model Merging and Guided Decoding for Controllable Multi-Objective Generation
Guofu Xie, Chen Zhang, Xiao Zhang +3
Adapting to diverse user needs at test time is a key challenge in controllable multi-objective generation. Existing methods are insufficient: merging-based approaches provide indir…
A Survey of Controllable Learning: Methods and Applications in Information Retrieval
Chenglei Shen, Xiao Zhang, Teng Shi +3
Controllability has become a crucial aspect of trustworthy machine learning, enabling learners to meet predefined targets and adapt dynamically at test time without requiring retra…
Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learning
Kepu Zhang, Guofu Xie, Weijie Yu +4
Legal mathematical reasoning is essential for applying large language models (LLMs) in high-stakes legal contexts, where outputs must be both mathematically accurate and procedural…
Bone Soups: A Seek-and-Soup Model Merging Approach for Controllable Multi-Objective Generation
Guofu Xie, Xiao Zhang, Ting Yao +1
User information needs are often highly diverse and varied. A key challenge in current research is how to achieve controllable multi-objective generation while enabling rapid adapt…