10 papers
Reliability-Prioritized Fine-Grained Generation in Multimodal Large
Xiaomeng Fan, Wei Wu, Yuwei Wu +9
Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theoretically show that generati…
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
Pengxiang Li, Zhi Gao, Bofei Zhang +8
Multimodal agents, which integrate a controller e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks.…
Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds
Wei Wu, Xiaomeng Fan, Yuwei Wu +4
Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods extract hierarchical features fro…
AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
Xintong Zhang, Xiaowen Zhang, Jingrong Wu +8
Adaptive multimodal reasoning has emerged as a promising frontier in Vision-Language Models (VLMs), aiming to dynamically modulate between tool-augmented visual reasoning and text…
GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
Chenrui Shi, Zedong Yu, Zhi Gao +7
Vision language models (VLMs) have advanced graphical user interface (GUI) task automation but still lag behind humans. We hypothesize this gap stems from missing core GUI knowledg…
Geometry-aware Distance Measure for Diverse Hierarchical Structures in Hyperbolic Spaces
Pengxiang Li, Yuwei Wu, Zhi Gao +5
Learning in hyperbolic spaces has attracted increasing attention due to its superior ability to model hierarchical structures of data. Most existing hyperbolic learning methods use…