14 citations · 14 across the 5 of their papers we have counts for
Showing 2025Show all
3 papers · 1 filter
cs.CV2025★ 14 cited
Qwen3-VL Technical Report
Shuai Bai, Yuxuan Cai, Ruizhe Chen +61
We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…
cs.CV2025
Enhancing Vision Foundation Models via Multimodal Continual Pre-Training
Yitong Chen, Lingchen Meng, Wujian Peng +4
Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this work, we enhance prevailing VFMs through multimodal training, allowi…
cs.CV2025
FOCUS: Towards Universal Foreground Segmentation
Zuyao You, Lingyu Kong, Lingchen Meng +1
Foreground segmentation is a fundamental task in computer vision, encompassing various subdivision tasks. Previous research has typically designed task-specific architectures for e…