activity
20242026
collaborators

8 papers

cs.CV2026

M-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

Jiatong Ma, Longteng Guo, Yuchen Liu +4

We present M-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multi…

cs.CV2025

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Daxiang Dong, Mingming Zheng, Dong Xu +32

We present Qianfan-VL, a series of multimodal large language models ranging from 3B to 70B parameters, achieving state-of-the-art performance through innovative domain enhancement…

cs.CV2025

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

Tongtian Yue, Longteng Guo, Yepeng Tang +4

Despite the impressive advancements of Large Vision-Language Models (LVLMs), existing approaches suffer from a fundamental bottleneck: inefficient visual-language integration. Curr…

cs.CV2025

OneDiff: A Generalist Model for Image Difference Captioning

Erdong Hu, Longteng Guo, Tongtian Yue +3

In computer vision, Image Difference Captioning (IDC) is crucial for accurately describing variations between closely related images. Traditional IDC methods often rely on speciali…

cs.CV2025

Image Difference Grounding with Natural Language

Wenxuan Wang, Zijia Zhao, Yisi Zhang +4

Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretat…

cs.CV2024

Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions

Wenxuan Wang, Yisi Zhang, Xingjian He +4

Visual grounding (VG) aims at locating the foreground entities that match the given natural language expressions. Previous datasets and methods for classic VG task mainly rely on t…