11 papers
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
Binghai Wang, Chenlong Zhang, Dayiheng Liu +10
A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this intuition is being inverted: as foundation models develop strong…
Temporal Motif-aware Graph Test-time Adaptation for OOD Blockchain Anomaly Detection
Runang He, Tongya Zheng, Huiling Peng +6
Ever-evolving transaction patterns have significantly hindered anomaly detection on emerging cryptocurrency blockchains due to the vast number of addresses and diverse anomalous be…
Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems
Zhezheng Hao, Tianfu Wang, Huanshuo Dong +7
LLM-based multi-agent systems (MAS) have emerged as an effective paradigm for complex and long-horizon tasks. However, in real-world tasks, MAS often exhibit various failures durin…
Battery-Sim-Agent: Leveraging LLM-Agent for Inverse Battery Parameter Estimation
Jiawei Chen, Xiaofan Gui, Shikai Fang +4
Parameterizing high-fidelity "digital twins" of batteries is a critical yet challenging inverse problem that hinders the pace of battery innovation. Prevailing methods formulate th…
MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models
Tianzhuo Yang, Zihan Shen, Zirui Mi +7
Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliabl…
LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks
Jiayong Wan, Jiawei Chen, Zhaoxia Yin +2
Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context reward hacking (ICRH), a phe…