12 papers
MuEvo: LLM-Driven Evolution of Multi-Heuristic Ensemble
Haoze Lv, Ning Lu, Shengcai Liu +2
Large language model-based automated heuristic design (LLM-AHD) has shown strong potential in discovering effective heuristics for combinatorial optimization problems. However, exi…
Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training
Zhiyuan Wang, Shengcai Liu, Jiahao Wu +5
Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution strategies (ES) enable memory-ef…
Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model
Jiahao Wu, Ning Lu, Shengcai Liu +6
Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks. While scaling rollouts can stabilize training and enhance perfor…
Policy and World Modeling Co-Training for Language Agents
Ning Lu, Baijiong Lin, Shengcai Liu +9
Reinforcement learning (RL) improves large language model (LLM) agents by teaching them which actions lead to high rewards, but provides little supervision on what those actions do…
AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design
Haoze Lv, Ning Lu, Ziang Zhou +2
Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent works show that large language models (L…
LLM-Driven Instance-Specific Heuristic Generation and Selection
Shaofeng Zhang, Shengcai Liu, Ning Lu +4
Combinatorial optimization problems are widely encountered in real-world applications. A critical research challenge lies in designing high-quality heuristic algorithms that effici…