7 papers
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar +6
Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agent…
Office Comprehension Benchmark
Firoz Shaik, Mateus Picanço Lima Gomes, Tanvir Aumi +17
We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.do…
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
Tri Cao, Yulin Chen, Hieu Cao +8
Web agents can autonomously complete online tasks by interacting with websites, but their exposure to open web environments makes them vulnerable to prompt injection attacks embedd…
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
Tri Cao, Khoi Le, Thong Nguyen +7
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure i…
REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning
Zhaoyuan Gu, Yipu Chen, Zimeng Chai +12
Humanoid loco-manipulation requires coordinated task-space motion planning with stable loco-manipulation command tracking under complex robot-environment dynamics and long-horizon…
VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents
Thong Q. Nguyen, Shubhang Desai, Raja Hasnain Anwar +3
Continual memory augmentation lets computer-using agents (CUAs) learn from prior interactions, but unvetted memories can encode domain-inappropriate or unsafe heuristics--spurious…