activity
20242026
most citedQwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

84 citations · 102 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV2026

SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation

Wei Tang, Xuejing Liu, Yanpeng Sun +1

The Segment Anything Model (SAM) excels at general image segmentation but has limited ability to understand natural language, which restricts its direct application in Referring Ex…

cs.CV202516 cited

Qwen3-VL Technical Report

Shuai Bai, Yuxuan Cai, Ruizhe Chen +61

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…

cs.CV2025

Revisiting Multimodal Positional Encoding in Vision-Language Models

Jie Huang, Xuejing Liu, Sibo Song +4

Multimodal position encoding is essential for vision-language models, yet there has been little systematic investigation into multimodal position encoding. We conduct a comprehensi…

cs.CL20252 cited

Qwen3-Omni Technical Report

Jin Xu, Zhifang Guo, Hangrui Hu +35

We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relat…

cs.CV2025

Qwen2.5-VL Technical Report

Shuai Bai, Keqin Chen, Xuejing Liu +24

We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative func…

cs.CV2024

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Zhibo Yang, Jun Tang, Zhaohai Li +9

Large Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what exten…