most citedQwen2.5-1M Technical Report

12 citations · 13 across the 33 of their papers we have counts for

collaborators

37 papers

cs.CV2026

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

Yijun Yang, Shenghe Zheng, Wenbo Li +8

Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue th…

cs.CV2026

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

Haoyu Wang, Songchun Zhang, Haoran Li +3

Action-conditioned video models require large-scale visual data paired with control signals that are temporally aligned with the resulting scene transitions. Such supervision is di…

cs.CV2026

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

Xionghao Wu, Yijun Yang, Shiyang Zhou +17

Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect…

cs.CV2026

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Nan Duan, Haoyang Huang, Weiyang Jin +13

Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain…

cs.CV2026

EchoWM: Open and Enterable Omnimodal World Models

Songchun Zhang, Yaowei Li, Junhao Zhuang +19

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music an…

cs.CV2026

HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models

Weilin Jin, Mingyu Wang, Wenbo Li +5

Although Multimodal Large Language Models have achieved strong performance across a wide range of vision-language tasks, they still suffer from hallucinations, where model outputs…