From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
Huatao Li, Xinwei Geng, Yuheng Wang +9
The paper presents DevicesWorld, a large executable benchmark of 6,140 tasks that require LLM‑based agents to operate across mobile, desktop, and IoT devices, and shows that curren…
cs.CL2026
FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection
Y. H. Zhou, Z. M. Ma, Y. J. Zhou +12
SMS fraud is increasingly cross-channel: a message directs the user to a webpage, and the final risk depends on how the SMS claim aligns with the page content and requested user ac…