34 papers
VTO: Visual Tool Orchestration for Video Anomaly Detection
Rui Wang, Yeteng Wu, Xianling Zhang +1
Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learning approaches are fundamental…
Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving
Mengshi Qi, Xiaoyang Bi, Xianlin Zhang +1
Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather conditions due to sensor perception degradation. These challenges arise fr…
CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving
Zhaohong Liu, Hao Ye, Xianlin Zhang +1
End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised Fine-Tuning (SFT) often suffers from reasoning hallucinations an…
A DVDrive Approach for doScenes Instructed Driving Challenge
Zijian Fu, Xiangyang Chu, Mengshi Qi +3
Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and…
Question-Aware Evidence Ledgers for Video Relational Reasoning
Yilin Ou, Mengshi Qi, Huadong Ma
The VRR-QA challenge evaluates visual relational reasoning in videos, where answers often depend on implicit spatial relations, event boundaries, target identity, and dialogue cont…
SGFormer++: Semantic Graph Transformer for Incremental 3D Scene Graph Generation
Mengshi Qi, Changsheng Lv, Zijian Fu +2
In this paper, we propose SGFormer++, a novel Semantic Graph Transformer for 3D scene graph generation (SGG), which aims to parse point cloud scenes into semantic structural graphs…