Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
PRO-CUA: Process-Reward Optimization for Computer Use Agents
Yifei He, Rui Yang, Hao Bai +2
Computer use agents (CUAs) have shown strong potential for automating complex digital workflows, yet their training remains constrained by costly live environment interaction and l…
cs.AI2025
ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning
Hanyang Chen, Mark Zhao, Rui Yang +15
Recent advances in embodied AI highlight the potential of vision language models (VLMs) as agents capable of perception, reasoning, and interaction in complex environments. However…
cs.AI2024
IWISDM: Assessing instruction following in multimodal models at scale
Xiaoxuan Lei, Lucas Gomez, Hao Yuan Bai +1
The ability to perform complex tasks from detailed instructions is a key to many remarkable achievements of our species. As humans, we are not only capable of performing a wide var…