8 papers
M-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
Jiatong Ma, Longteng Guo, Yuchen Liu +4
We present M-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multi…
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
Daxiang Dong, Mingming Zheng, Dong Xu +32
We present Qianfan-VL, a series of multimodal large language models ranging from 3B to 70B parameters, achieving state-of-the-art performance through innovative domain enhancement…
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
Tongtian Yue, Longteng Guo, Yepeng Tang +4
Despite the impressive advancements of Large Vision-Language Models (LVLMs), existing approaches suffer from a fundamental bottleneck: inefficient visual-language integration. Curr…
OneDiff: A Generalist Model for Image Difference Captioning
Erdong Hu, Longteng Guo, Tongtian Yue +3
In computer vision, Image Difference Captioning (IDC) is crucial for accurately describing variations between closely related images. Traditional IDC methods often rely on speciali…
Image Difference Grounding with Natural Language
Wenxuan Wang, Zijia Zhao, Yisi Zhang +4
Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretat…
Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions
Wenxuan Wang, Yisi Zhang, Xingjian He +4
Visual grounding (VG) aims at locating the foreground entities that match the given natural language expressions. Previous datasets and methods for classic VG task mainly rely on t…