1 citations · 1 across the 5 of their papers we have counts for
13 papers · 1 filter
DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection
Yuyang Hong, Jinhui Guo, Jiaqi Gu +6
Visual instruction tuning is crucial for advancing the vision-language alignment and instruction-following capabilities of Vision-Language Models (VLMs). However, identifying optim…
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
Jintao Tong, Jiaqi Gu, Yujing Lou +5
While Multimodal Large Language Models (MLLMs) excel at visual understanding tasks through text reasoning, they often fall short in scenarios requiring visual imagination. Unlike c…
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
Yuyang Hong, Jiaqi Gu, Qi Yang +6
Knowledge-based visual question answering (KB-VQA) requires visual language models (VLMs) to integrate visual understanding with external knowledge retrieval. Although retrieval-au…
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
Pingyi Chen, Yujing Lou, Shen Cao +6
While vision language models (VLMs) excel in 2D semantic visual understanding, their ability to quantitatively reason about 3D spatial relationships remains under-explored, due to…
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
Shixiong Xu, Chenghao Zhang, Lubin Fan +5
Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained s…
Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger
Qi Yang, Chenghao Zhang, Lubin Fan +3
Recent advancements in Large Vision Language Models (LVLMs) have significantly improved performance in Visual Question Answering (VQA) tasks through multimodal Retrieval-Augmented…