10 papers
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
A governance horizon for ethical-use constraints in open-weight AI models
Weiwei Xu, Hengzhi Ye, Haoran Ye +3
Ethical constraints on open-weight AI models are both a reflection of societal concerns and a foundation for AI governance policy. They are expected to propagate to downstream deri…
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
Haonan Dong, Qiguan Feng, Kehan Jiang +3
Autonomous agents have rapidly matured as task executors and seen widespread deployment via harnesses such as OpenClaw. Safety concerns have rightly drawn growing research attentio…
NeuReasoner: Towards Explainable, Controllable, and Unified Reasoning via Mixture-of-Neurons
Haonan Dong, Kehan Jiang, Haoran Ye +3
Large Reasoning Models (LRMs) have recently achieved remarkable success in complex reasoning tasks. However, closer scrutiny reveals persistent failure modes compromising performan…
Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement
Haoran Ye, Jing Jin, Yuhang Xie +2
The advancement of large language models (LLMs) has outpaced traditional evaluation methodologies. This progress presents novel challenges, such as measuring human-like psychologic…
VRPAgent: LLM-Driven Discovery of Heuristic Operators for Vehicle Routing Problems
André Hottung, Federico Berto, Chuanbo Hua +9
Designing high-performing heuristics for vehicle routing problems (VRPs) is a complex task that requires both intuition and deep domain knowledge. Large language model (LLM)-based…