2 papers
cs.AI2026
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Jingbo Zhou, Yusai Zhao, Qi Bao +12
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents…
cs.CV2026
Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents
Jia-Chen Zhang, Ze-Yu Zhang, Kai-Wei Zhang
Computer-use agents (CUAs), which empower large language models to autonomously operate operating systems and the web, are increasingly vulnerable to indirect prompt injection atta…