6 papers
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
Shengyi Wang, Niantong Li, Guangzheng Hu +27
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM…
Discovering Millions of Interpretable Features with Sparse Autoencoders
XinYang He, Wei Wang, Bing Zhao +5
Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs…
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
Niantong Li, Guangzheng Hu, Weixu Qiao +35
Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no…
Qwen-Image-2.0 Technical Report
Bing Zhao, Chenfei Wu, Deqing Li +72
We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite rece…
Logics-Parsing-Omni Technical Report
Xin An, Jingyi Cai, Xiangyang Chen +22
Addressing the challenges of fragmented task definitions and the heterogeneity of unstructured data in multimodal parsing, this paper proposes the Omni Parsing framework. This fram…
Agile Deliberation: Concept Deliberation for Subjective Visual Classification
Leijie Wang, Otilia Stretcu, Wei Qiao +7
From content moderation to content curation, applications requiring vision classifiers for visual concepts are rapidly expanding. Existing human-in-the-loop approaches typically as…