16 citations · 20 across the 8 of their papers we have counts for
13 papers · 1 filter
Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ
Zongyun Zhang, Jiacheng Ruan, Xian Gao +5
Although multimodal large language models (MLLMs) have shown substantial potential in visual understanding and graphic code generation, editing scientific figures through code pres…
VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
Yuqi Zhang, Cheng Chen, Yuyu Guo +6
Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions eit…
Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification
Wujian Peng, Lingchen Meng, Yuxuan Cai +7
Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokeni…
Qwen3-VL Technical Report
Shuai Bai, Yuxuan Cai, Ruizhe Chen +61
We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…
Enhancing Vision Foundation Models via Multimodal Continual Pre-Training
Yitong Chen, Lingchen Meng, Wujian Peng +4
Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this work, we enhance prevailing VFMs through multimodal training, allowi…
FOCUS: Towards Universal Foreground Segmentation
Zuyao You, Lingyu Kong, Lingchen Meng +1
Foreground segmentation is a fundamental task in computer vision, encompassing various subdivision tasks. Previous research has typically designed task-specific architectures for e…