activity
20242026
collaborators

17 papers

cs.CV2026

FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction

Dubing Chen, Huan Zheng, Tianyi Yan +5

Vision-based 3D occupancy prediction fundamentally relies on the 2D-to-3D view transformation. Current paradigms predominantly utilize explicit physical projection, which artificia…

cs.CV2026

Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs

Huan Zheng, Yucheng Zhou, Tianyi Yan +6

Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in medical image analysis. However, their application in gastrointestinal endoscopy is currently hin…

cs.CV2026

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

Jie Zhang, Xiaoyue Chen, Anzhe Chen +36

We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically ground…

cs.CV2026

CausalDrive: Real-time Causal World Models for Autonomous Driving

Tianyi Yan, Huan Zheng, Dubing Chen +10

World models have emerged as a promising paradigm for scaling autonomous driving (AD) data, yet existing video generative models fall short as interactive simulators. Layout-condit…

cs.CV2026

Is Your Driving World Model an All-Around Player?

Lingdong Kong, Ao Liang, Tianyi Yan +20

Today's driving world models can generate remarkably realistic dash-cam videos, yet no single model excels universally. Some generate photorealistic textures but violate basic phys…

cs.CV2026

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)

Wei Luo, Yiting Lu, Xin Li +32

This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and 4D generation settings. The…