collaborators

7 papers

cs.AI2026

AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees

Yuankai Li, Tinghui Zhu, Ha Min Son +3

Large Multimodal Models (LMMs) have recently emerged as promising backbones for GUI-agent models, where high-resolution GUI screenshots are introduced to the prompts at each iterat…

cs.CV2026

Video Models Can Reason with Verifiable Rewards

Tinghui Zhu, Sheng Zhang, James Y. Huang +5

Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable re…

cs.CV2026

When Vision Speaks for Sound

Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu +6

Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual cues to infer or hallucinate…

cs.CV2026

Learning Adaptive Reasoning Paths for Efficient Visual Reasoning

Yixu Huang, Tinghui Zhu, Muhao Chen

Visual reasoning models (VRMs) have recently shown strong cross-modal reasoning capabilities by integrating visual perception with language reasoning. However, they often suffer fr…

cs.AI2025

OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning

Boyu Zhu, Xiaofei Wen, Wenjie Jacky Mo +4

Omni-modal Large Language Models (OLLMs) that process text, images, videos, and audio introduce new challenges for safety and value guardrails in human-AI interaction. Prior guardr…

cs.CL2025

Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration

James Y. Huang, Sheng Zhang, Qianchu Liu +5

Large Language Models (LLMs) have demonstrated remarkable capabilities in challenging, knowledge-intensive reasoning tasks. However, extending LLMs to perceive and reason over a ne…