activity
20242026
collaborators

34 papers

cs.CV2026

VTO: Visual Tool Orchestration for Video Anomaly Detection

Rui Wang, Yeteng Wu, Xianling Zhang +1

Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learning approaches are fundamental…

cs.CV2026

Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving

Mengshi Qi, Xiaoyang Bi, Xianlin Zhang +1

Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather conditions due to sensor perception degradation. These challenges arise fr…

cs.CV2026

CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving

Zhaohong Liu, Hao Ye, Xianlin Zhang +1

End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised Fine-Tuning (SFT) often suffers from reasoning hallucinations an…

cs.CV2026

A DVDrive Approach for doScenes Instructed Driving Challenge

Zijian Fu, Xiangyang Chu, Mengshi Qi +3

Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and…

cs.CV2026

Question-Aware Evidence Ledgers for Video Relational Reasoning

Yilin Ou, Mengshi Qi, Huadong Ma

The VRR-QA challenge evaluates visual relational reasoning in videos, where answers often depend on implicit spatial relations, event boundaries, target identity, and dialogue cont…

cs.CV2026

SGFormer++: Semantic Graph Transformer for Incremental 3D Scene Graph Generation

Mengshi Qi, Changsheng Lv, Zijian Fu +2

In this paper, we propose SGFormer++, a novel Semantic Graph Transformer for 3D scene graph generation (SGG), which aims to parse point cloud scenes into semantic structural graphs…