3 papers
cs.LG2026
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Zhi Zheng, Rongsheng Chen, Yunpeng Ba +3
Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rew…
cs.AI2026
Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging
Yu Gu, Zhi Zheng, Yunpeng Ba +3
Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However, directly applying ES to bil…
math.OC2026
Survey on Neural Routing Solvers
Yunpeng Ba, Xi Lin, Changliang Zhou +7
Neural routing solvers (NRSs) that leverage deep learning to tackle vehicle routing problems have demonstrated notable potential for practical applications. By learning implicit he…