3 papers
cs.CL2026
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
Huatao Li, Xinwei Geng, Yuheng Wang +9
The paper presents DevicesWorld, a large executable benchmark of 6,140 tasks that require LLM‑based agents to operate across mobile, desktop, and IoT devices, and shows that curren…
cs.AI2026
MagicAgent: Towards Generalized Agent Planning
Xuhui Ren, Shaokang Dong, Chen Yang +21
The evolution of Large Language Models (LLMs) from passive text processors to autonomous agents has established planning as a core component of modern intelligence. However, achiev…
cs.AI2026
Know What You Know: Metacognitive Entropy Calibration for Verifiable RL Reasoning
Qiannian Zhao, Chen Yang, Jinhao Jing +5
Large reasoning models (LRMs) have emerged as a powerful paradigm for solving complex real-world tasks. In practice, these models are predominantly trained via Reinforcement Learni…