6 papers
OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation
Haochen Yang, Ke Zhao, Mengyuan Ma +3
Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient paradigm for automated optimiza…
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
Hao Wang, Haocheng Yang, Licheng Pan +7
Reward modeling represents a long-standing challenge in reinforcement learning from human feedback (RLHF) for aligning language models. Current reward modeling is heavily contingen…
Reducing Belief Deviation in Reinforcement Learning for Active Reasoning
Deyu Zou, Yongqiang Chen, Jianxiang Wang +5
Active reasoning requires large language model (LLM) agents to interact with external sources and strategically gather information to solve problems in multiple turns. Central to t…
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
Hao Chen, Zhexin Hu, Jiajun Chai +7
Training LLMs to invoke tools and leverage retrieved information necessitates high-quality, diverse data. However, existing pipelines for synthetic data generation often rely on te…
Adaptive Selection of Symbolic Languages for Improving LLM Logical Reasoning
Xiangyu Wang, Haocheng Yang, Fengxiang Cheng +1
Large Language Models (LLMs) still struggle with complex logical reasoning. While previous works achieve remarkable improvements, their performance is highly dependent on the corre…
Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling
Han Qi, Haochen Yang, Qiaosheng Zhang +1
We study the problem of reinforcement learning from human feedback (RLHF), a critical problem in training large language models, from a theoretical perspective. Our main contributi…