activity
20192026
most citedBEVerse: Unified Perception and Prediction in Birds-Eye-View for Vision-Centric Autonomous Driving

84 citations · 104 across the 44 of their papers we have counts for

collaborators

47 papers

cs.CV2026

SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

Shing Ho J. Lin, Wenzhao Zheng, Dong Zhuo +3

Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding remains a fundamental challenge…

cs.CV2026

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection

Yanran Zhang, Wenzhao Zheng, Yifei Li +5

In recent years, significant progress has been made in both image generation and generated image detection. Despite their rapid, yet largely independent, development, these two fie…

cs.CV2026

Vega: Learning to Drive with Natural Language Instructions

Sicheng Zuo, Yuxuan Li, Wenzhao Zheng +3

Vision-language-action models have reshaped autonomous driving to incorporate languages into the decision-making process. However, most existing pipelines only utilize the language…

cs.CV2026

DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding

Dong Zhuo, Wenzhao Zheng, Sicheng Zuo +4

With the growing adoption of vision-language-action models and world models in autonomous driving systems, scalable image tokenization becomes crucial as the interface for the visu…

cs.CV2026

Astra: General Interactive World Model with Autoregressive Denoising

Yixuan Zhu, Jiaqi Feng, Wenzhao Zheng +5

Recent advances in diffusion transformers have empowered video generation models to generate high-quality video clips from texts or images. However, world models with the ability t…

cs.CV2026

Moaw: Unleashing Motion Awareness for Video Diffusion Models

Tianqi Zhang, Ziyi Wang, Wenzhao Zheng +5

Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks suc…