works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CL2026

DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

Huatao Li, Xinwei Geng, Yuheng Wang +9

The paper presents DevicesWorld, a large executable benchmark of 6,140 tasks that require LLM‑based agents to operate across mobile, desktop, and IoT devices, and shows that curren…

cs.CL2026

Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems

Shu Yao, Yuhua Luo, Qian Long +7

Real-world computer-use tasks often span multiple applications and devices, requiring agents to coordinate heterogeneous environments under dynamic runtime failures. Existing multi…

cs.AI2026

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

Qiran Zhang, Yuheng Wang, Runde Yang +9

Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models…

cs.CL2026

Towards a Science of Collective AI: LLM-based Multi-Agent Systems Need a Transition from Blind Trial-and-Error to Rigorous Science

Jingru Fan, Dewen Liu, Yufan Dang +15

Recent advancements in Large Language Models (LLMs) have greatly extended the capabilities of Multi-Agent Systems (MAS), demonstrating significant effectiveness across a wide range…

cs.LG2025

RLPR: Extrapolating RLVR to General Domains without Verifiers

Tianyu Yu, Bo Ji, Shouli Wang +9

Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates promising potential in advancing the reasoning capabilities of LLMs. However, its success remains largely confine…