6 papers
WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
Sanjari Srivastava, Gang Li, Cheng Chang +8
Training web agents to navigate complex, real-world websites requires them to master - short-horizon interactions on multiple UI components (e.g., choosing the…
Learning Efficient Guardrails for Compliance
Xiaofei Wen, Wenjie Jacky Mo, Yanan Xie +2
Autonomous web agents are increasingly deployed for long-horizon tasks, yet their ability to adhere to real-world policies remains critically underexplored compared to standard saf…
REC-RL: Referring expression counting via Gaussian and range-based reward optimization
Hui Liu, Yunlai Teng, Kunlong Bai +4
Referring expression counting (REC) is an intention-driven task that requires context-aware visual reasoning. While recent vision-language models incorporate language for visual un…
When Vision Speaks for Sound
Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu +6
Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual cues to infer or hallucinate…
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
Boyu Zhu, Xiaofei Wen, Wenjie Jacky Mo +4
Omni-modal Large Language Models (OLLMs) that process text, images, videos, and audio introduce new challenges for safety and value guardrails in human-AI interaction. Prior guardr…
Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
Yu Gu, Kai Zhang, Yuting Ning +9
Language agents based on large language models (LLMs) have demonstrated great promise in automating web-based tasks. Recent work has shown that incorporating advanced planning algo…