Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training
Zhiyuan Wang, Shengcai Liu, Jiahao Wu +5
Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution strategies (ES) enable memory-ef…
cs.AI2026
AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design
Haoze Lv, Ning Lu, Ziang Zhou +2
Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent works show that large language models (L…
cs.AI2025
Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
Zhangying Feng, Qianglong Chen, Ning Lu +6
The development of reasoning capabilities represents a critical frontier in large language models (LLMs) research, where reinforcement learning (RL) and process reward models (PRMs…