5 papers · 1 filter
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
A governance horizon for ethical-use constraints in open-weight AI models
Weiwei Xu, Hengzhi Ye, Haoran Ye +3
Ethical constraints on open-weight AI models are both a reflection of societal concerns and a foundation for AI governance policy. They are expected to propagate to downstream deri…
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
Haonan Dong, Qiguan Feng, Kehan Jiang +3
Autonomous agents have rapidly matured as task executors and seen widespread deployment via harnesses such as OpenClaw. Safety concerns have rightly drawn growing research attentio…
VRPAgent: LLM-Driven Discovery of Heuristic Operators for Vehicle Routing Problems
André Hottung, Federico Berto, Chuanbo Hua +9
Designing high-performing heuristics for vehicle routing problems (VRPs) is a complex task that requires both intuition and deep domain knowledge. Large language model (LLM)-based…
GLOP: Learning Global Partition and Local Construction for Solving Large-scale Routing Problems in Real-time
Haoran Ye, Jiarui Wang, Helan Liang +3
The recent end-to-end neural solvers have shown promise for small-scale routing problems but suffered from limited real-time scaling-up performance. This paper proposes GLOP (Globa…