activity
20242026
collaborators

10 papers

cs.CV2026

Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation

Jintao Sun, Gangyi Ding, Donglin Di +2

Vision-Language Models have achieved strong progress in ground-view visual understanding, yet they remain brittle in high-altitude Unmanned Aerial Vehicle scenes, where objects are…

cs.CV2026

Uncertainty-Aware Trajectory Prediction: A Unified Framework Harnessing Positional and Semantic Uncertainties

Jintao Sun, Hu Zhang, Gangyi Ding +1

Trajectory prediction seeks to forecast the future motion of dynamic entities, such as vehicles and pedestrians, given a temporal horizon of historical movement data and environmen…

cs.CV2026

Echo Planning for Autonomous Driving: From Current Observations to Future Trajectories and Back

Jintao Sun, Hu Zhang, Gangyi Ding +1

Modern end-to-end autonomous driving systems suffer from a critical limitation: their planners lack mechanisms to enforce temporal consistency between predicted trajectories and ev…

cs.RO2025

Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning

Xun Li, Rodrigo Santa Cruz, Mingze Xi +9

To enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and intera…

cs.CV2025

AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search

Hao Ju, Hu Zhang, Zhedong Zheng

With growing public safety demands, text-based person anomaly search has emerged as a critical task, aiming to retrieve individuals with abnormal behaviors via natural language des…

cs.CV2025

Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention

Yiwen Chen, Zhihao Li, Yikai Wang +4

Recent advances in sparse voxel representations have significantly improved the quality of 3D content generation, enabling high-resolution modeling with fine-grained geometry. Howe…