3 citations · 3 across the 4 of their papers we have counts for
4 papers
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
Zixuan Dong, Baoyun Peng, Yufei Wang +4
Human video comprehension demonstrates dynamic coordination between reasoning and visual attention, adaptively focusing on query-relevant details. However, current long-form video…
LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model
Xiaodong Wang, Zhirong Wu, Peixi Peng
Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error ac…
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
Zehuan Wu, Jingcheng Ni, Xiaodong Wang +5
Generative models have significantly improved the generation and prediction quality on either camera images or LiDAR point clouds for autonomous driving. However, a real-world auto…
Learning Invariant Representation with Consistency and Diversity for Semi-supervised Source Hypothesis Transfer
Xiaodong Wang, Junbao Zhuo, Shuhao Cui +1
Semi-supervised domain adaptation (SSDA) aims to solve tasks in target domain by utilizing transferable information learned from the available source domain and a few labeled targe…