collaborators

6 papers

cs.AI2026

OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation

Haochen Yang, Ke Zhao, Mengyuan Ma +3

Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient paradigm for automated optimiza…

cs.CL2026

ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment

Hao Wang, Haocheng Yang, Licheng Pan +7

Reward modeling represents a long-standing challenge in reinforcement learning from human feedback (RLHF) for aligning language models. Current reward modeling is heavily contingen…

cs.AI2026

Reducing Belief Deviation in Reinforcement Learning for Active Reasoning

Deyu Zou, Yongqiang Chen, Jianxiang Wang +5

Active reasoning requires large language model (LLM) agents to interact with external sources and strategically gather information to solve problems in multiple turns. Central to t…

cs.AI2025

ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs

Hao Chen, Zhexin Hu, Jiajun Chai +7

Training LLMs to invoke tools and leverage retrieved information necessitates high-quality, diverse data. However, existing pipelines for synthetic data generation often rely on te…

cs.AI2025

Adaptive Selection of Symbolic Languages for Improving LLM Logical Reasoning

Xiangyu Wang, Haocheng Yang, Fengxiang Cheng +1

Large Language Models (LLMs) still struggle with complex logical reasoning. While previous works achieve remarkable improvements, their performance is highly dependent on the corre…

cs.LG2025

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling

Han Qi, Haochen Yang, Qiaosheng Zhang +1

We study the problem of reinforcement learning from human feedback (RLHF), a critical problem in training large language models, from a theoretical perspective. Our main contributi…