activity
20212026
most citedOPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

21 citations · 31 across the 7 of their papers we have counts for

collaborators

9 papers

cs.CV2026

M-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

Jiatong Ma, Longteng Guo, Yuchen Liu +4

We present M-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multi…

cs.CV2025

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Daxiang Dong, Mingming Zheng, Dong Xu +32

We present Qianfan-VL, a series of multimodal large language models ranging from 3B to 70B parameters, achieving state-of-the-art performance through innovative domain enhancement…

cs.CV2025

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

Tongtian Yue, Longteng Guo, Yepeng Tang +4

Despite the impressive advancements of Large Vision-Language Models (LVLMs), existing approaches suffer from a fundamental bottleneck: inefficient visual-language integration. Curr…

cs.CV2025

Image Difference Grounding with Natural Language

Wenxuan Wang, Zijia Zhao, Yisi Zhang +4

Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretat…

cs.CV2024

OneDiff: A Generalist Model for Image Difference Captioning

Erdong Hu, Longteng Guo, Tongtian Yue +3

In computer vision, Image Difference Captioning (IDC) is crucial for accurately describing variations between closely related images. Traditional IDC methods often rely on speciali…

cs.CV202410 cited

VL-Mamba: Exploring State Space Models for Multimodal Learning

Yanyuan Qiao, Zheng Yu, Longteng Guo +5

Multimodal large language models (MLLMs) have attracted widespread interest and have rich applications. However, the inherent attention mechanism in its Transformer structure requi…