most citedQwen-Image Technical Report

2 citations · 2 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing

Zijian Kan, Wei Wang, Long Luo +6

Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot ad…

cs.CV2026

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning

Junzhe Zhang, Huixuan Zhang, Guirong Wang +5

With the rapid advancement of multimodal large language models (MLLMs), models have demonstrated increasingly powerful multimodal capabilities. However, whether MLLMs trained throu…

cs.CV2026

Qwen-Image-VAE-2.0 Technical Report

Zekai Zhang, Deqing Li, Kuan Cao +27

We present Qwen-Image-VAE-2.0, a suite of high-compression Variational Autoencoders (VAEs) that achieve significant advances in both reconstruction fidelity and diffusability. To a…

cs.CV2026

Qwen-Image-2.0 Technical Report

Bing Zhao, Chenfei Wu, Deqing Li +72

We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite rece…

cs.CV2025

Logics-Parsing Technical Report

Xiangyang Chen, Shuzhao Li, Xiuwen Zhu +7

Recent advances in Large Vision-Language models (LVLM) have spurred significant progress in document parsing task. Compared to traditional pipeline-based methods, end-to-end paradi…

cs.CV20252 cited

Qwen-Image Technical Report

Chenfei Wu, Jiahao Li, Jingren Zhou +36

We present Qwen-Image, an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. To address th…