12 papers
FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction
Dubing Chen, Huan Zheng, Tianyi Yan +5
Vision-based 3D occupancy prediction fundamentally relies on the 2D-to-3D view transformation. Current paradigms predominantly utilize explicit physical projection, which artificia…
CausalDrive: Real-time Causal World Models for Autonomous Driving
Tianyi Yan, Huan Zheng, Dubing Chen +10
World models have emerged as a promising paradigm for scaling autonomous driving (AD) data, yet existing video generative models fall short as interactive simulators. Layout-condit…
Multimodal Large Language Models for Multi-Subject In-Context Image Generation
Yucheng Zhou, Dubing Chen, Huan Zheng +1
Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains…
EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents
Xuan Dong, Huanyang Zheng, Tianhao Niu +7
Scientific research follows multi-turn, multi-step workflows that require proactively searching the literature, consulting figures and tables, and integrating evidence across paper…
Towards Geometry-Aware and Motion-Guided Video Human Mesh Recovery
Hongjun Chen, Huan Zheng, Wencheng Han +1
Existing video-based 3D Human Mesh Recovery (HMR) methods often produce physically implausible results, stemming from their reliance on flawed intermediate 3D pose anchors and thei…
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
Huan Zheng, Yucheng Zhou, Tianyi Yan +9
While end-to-end autonomous driving has achieved remarkable progress in geometric control, current systems remain constrained by a command-following paradigm that relies on simple…