3 papers
cs.CV2026
Reliability-Prioritized Fine-Grained Generation in Multimodal Large
Xiaomeng Fan, Wei Wu, Yuwei Wu +9
Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theoretically show that generati…
cs.CV2026
Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds
Wei Wu, Xiaomeng Fan, Yuwei Wu +4
Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods extract hierarchical features fro…
cs.CV2025
MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization
Zhendong Xiao, Wu Wei, Shujie Ji +2
Camera relocalization, a cornerstone capability of modern computer vision, accurately determines a camera's position and orientation (6-DoF) from images and is essential for applic…