10 papers
DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs
Anqi Li, Jie Zhang, Zhongqi Wang +4
While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predomi…
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
Niantong Li, Guangzheng Hu, Weixu Qiao +35
Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no…
Qwen-Image-VAE-2.0 Technical Report
Zekai Zhang, Deqing Li, Kuan Cao +27
We present Qwen-Image-VAE-2.0, a suite of high-compression Variational Autoencoders (VAEs) that achieve significant advances in both reconstruction fidelity and diffusability. To a…
Qwen-Image-2.0 Technical Report
Bing Zhao, Chenfei Wu, Deqing Li +72
We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite rece…
ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration
Bei Yan, Yuecong Min, Jie Zhang +2
Large Vision-Language Models (LVLMs) frequently suffer from severe hallucination issues. Existing mitigation strategies predominantly rely on isolated, single-step states to enhanc…
A Survey of Multimodal Hallucination Evaluation and Detection
Zhiyuan Chen, Yuecong Min, Jie Zhang +4
Multi-modal Large Language Models (MLLMs) have emerged as a powerful paradigm for integrating visual and textual information, supporting a wide range of multi-modal tasks. However,…