104 citations · 300 across the 6 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024★ 84 cited
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Peng Wang, Shuai Bai, Sinan Tan +16
We present the Qwen2-VL Series, an advanced upgrade of the previous Qwen-VL models that redefines the conventional predetermined-resolution approach in visual processing. Qwen2-VL…
cs.CV2024
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
Liang Chen, Sinan Tan, Zefan Cai +7
This work tackles the information loss bottleneck of vector-quantization (VQ) autoregressive image generation by introducing a novel model architecture called the 2-Dimensional Aut…
cs.CV2022★ 3 cited
OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models
Jinze Bai, Rui Men, Hao Yang +15
Generalist models, which are capable of performing diverse multi-modal tasks in a task-agnostic way within a single model, have been explored recently. Being, hopefully, an alterna…