12 papers
LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
Zhou Tao, Fang Zhang, Zewen Ding +5
Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify thi…
Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation
Kyle Gao, Joel Cumming, Jonathan Li +2
We present an LLM-driven framework for retrieving remote sensing data from cloud-based geospatial catalogues using natural language queries. The system converts user intent into st…
Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning
Shida Wang, YongXiang Hua, Zhou Tao +2
Multimodal Large Language Models have demonstrated remarkable capabilities in video understanding, yet face prohibitive computational costs and performance degradation from ''conte…
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
Xiaokun Sun, Yubo Wang, Haoyu Cao +1
Recently, Multimodal Large Language Models (MLLMs) have demonstrated significant potential in complex visual tasks through the integration of Chain-of-Thought (CoT) reasoning. Howe…
DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
Zhou Tao, Shida Wang, Yongxiang Hua +2
Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning…
Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
Yongxiang Hua, Haoyu Cao, Zhou Tao +4
Sparse Mixture of Experts (sMoE) has become a pivotal approach for scaling large vision-language models, offering substantial capacity while maintaining computational efficiency th…