activity
20242026
collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling

Yuqi Zhang, Cheng Chen, Yuyu Guo +6

Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions eit…

cs.CV2026

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

Wujian Peng, Lingchen Meng, Yuxuan Cai +7

Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokeni…

cs.CV2026

INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning

Wujian Peng, Lingchen Meng, Yitong Chen +7

Large Multimodal Models (LMMs) have made significant breakthroughs with the advancement of instruction tuning. However, while existing models can understand images and videos at a…

cs.CV202514 cited

Qwen3-VL Technical Report

Shuai Bai, Yuxuan Cai, Ruizhe Chen +61

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…

cs.CV2025

Enhancing Vision Foundation Models via Multimodal Continual Pre-Training

Yitong Chen, Lingchen Meng, Wujian Peng +4

Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this work, we enhance prevailing VFMs through multimodal training, allowi…

cs.CV2025

FOCUS: Towards Universal Foreground Segmentation

Zuyao You, Lingyu Kong, Lingchen Meng +1

Foreground segmentation is a fundamental task in computer vision, encompassing various subdivision tasks. Previous research has typically designed task-specific architectures for e…