7 papers
AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees
Yuankai Li, Tinghui Zhu, Ha Min Son +3
Large Multimodal Models (LMMs) have recently emerged as promising backbones for GUI-agent models, where high-resolution GUI screenshots are introduced to the prompts at each iterat…
Video Models Can Reason with Verifiable Rewards
Tinghui Zhu, Sheng Zhang, James Y. Huang +5
Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable re…
When Vision Speaks for Sound
Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu +6
Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual cues to infer or hallucinate…
Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
Yixu Huang, Tinghui Zhu, Muhao Chen
Visual reasoning models (VRMs) have recently shown strong cross-modal reasoning capabilities by integrating visual perception with language reasoning. However, they often suffer fr…
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
Boyu Zhu, Xiaofei Wen, Wenjie Jacky Mo +4
Omni-modal Large Language Models (OLLMs) that process text, images, videos, and audio introduce new challenges for safety and value guardrails in human-AI interaction. Prior guardr…
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
James Y. Huang, Sheng Zhang, Qianchu Liu +5
Large Language Models (LLMs) have demonstrated remarkable capabilities in challenging, knowledge-intensive reasoning tasks. However, extending LLMs to perceive and reason over a ne…