1 citations · 1 across the 3 of their papers we have counts for
4 papers
ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
Joy Chen, Alejandro Castillejo Munoz, Pierluca D'Oro +3
Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on behalf of users, yet no benchmark comprehensively evaluates whether they can saf…
Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance
Kaustav Kundu, Ritvik Shrivastava, Maxim Arap +13
We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding \textit{when} to interrupt, and \…
The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes
Redacted by arXiv
This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
Yuxuan Sun, Manchen Wang, Shengyi Qian +18
AI agents capable of controlling user interfaces have the potential to transform human interaction with digital devices. To accelerate this transformation, two fundamental building…