most citedQwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

84 citations · 86 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CV20252 cited

Qwen-Image Technical Report

Chenfei Wu, Jiahao Li, Jingren Zhou +36

We present Qwen-Image, an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. To address th…

cs.IR2025

DAS: Dual-Aligned Semantic IDs Empowered Industrial Recommender System

Wencai Ye, Mingjie Sun, Shaoyun Shi +3

Semantic IDs are discrete identifiers generated by quantizing the Multi-modal Large Language Models (MLLMs) embeddings, enabling efficient multi-modal content integration in recomm…

cs.CL2025

InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes under Herd Behavior

Huisheng Wang, Zhuoshi Pan, Hangjing Zhang +3

Aligning Large Language Models (LLMs) with investor decision-making processes under herd behavior is a critical challenge in behavioral finance, which grapples with a fundamental l…

cs.CV2025

Qwen2.5-VL Technical Report

Shuai Bai, Keqin Chen, Xuejing Liu +24

We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative func…

cs.CV202484 cited

Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Peng Wang, Shuai Bai, Sinan Tan +16

We present the Qwen2-VL Series, an advanced upgrade of the previous Qwen-VL models that redefines the conventional predetermined-resolution approach in visual processing. Qwen2-VL…