5 papers
Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos
Dingkun Wei, Zehong Shen, Yan Xia +3
Human motion recovered from monocular videos often appears overly smooth or dynamically inconsistent, even when joint positions are numerically accurate. We observe that this limit…
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
Peng Zhang, Guanghao Zhang, Wanggui He +10
Recent video multimodal large language models (MLLMs) increasingly couple step-by-step reasoning with on-demand visual evidence retrieval, allowing models to revisit relevant video…
STDDN: A Physics-Guided Deep Learning Framework for Crowd Simulation
Zijin Liu, Xu Geng, Wenshuai Xu +3
Accurate crowd simulation is crucial for public safety management, emergency evacuation planning, and intelligent transportation systems. However, existing methods, which typically…
Energy-Aware Imitation Learning for Steering Prediction Using Events and Frames
Hu Cao, Jiong Liu, Xingzhuo Yan +5
In autonomous driving, relying solely on frame-based cameras can lead to inaccuracies caused by factors like long exposure times, high-speed motion, and challenging lighting condit…
SparseAlign: A Fully Sparse Framework for Cooperative Object Detection
Yunshuang Yuan, Yan Xia, Daniel Cremers +1
Cooperative perception can increase the view field and decrease the occlusion of an ego vehicle, hence improving the perception performance and safety of autonomous driving. Despit…