10 papers
Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
Jintao Sun, Gangyi Ding, Donglin Di +2
Vision-Language Models have achieved strong progress in ground-view visual understanding, yet they remain brittle in high-altitude Unmanned Aerial Vehicle scenes, where objects are…
Uncertainty-Aware Trajectory Prediction: A Unified Framework Harnessing Positional and Semantic Uncertainties
Jintao Sun, Hu Zhang, Gangyi Ding +1
Trajectory prediction seeks to forecast the future motion of dynamic entities, such as vehicles and pedestrians, given a temporal horizon of historical movement data and environmen…
Echo Planning for Autonomous Driving: From Current Observations to Future Trajectories and Back
Jintao Sun, Hu Zhang, Gangyi Ding +1
Modern end-to-end autonomous driving systems suffer from a critical limitation: their planners lack mechanisms to enforce temporal consistency between predicted trajectories and ev…
Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
Xun Li, Rodrigo Santa Cruz, Mingze Xi +9
To enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and intera…
AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
Hao Ju, Hu Zhang, Zhedong Zheng
With growing public safety demands, text-based person anomaly search has emerged as a critical task, aiming to retrieve individuals with abnormal behaviors via natural language des…
Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
Yiwen Chen, Zhihao Li, Yikai Wang +4
Recent advances in sparse voxel representations have significantly improved the quality of 3D content generation, enabling high-resolution modeling with fine-grained geometry. Howe…