collaborators

9 papers

cs.CV2026

VTO: Visual Tool Orchestration for Video Anomaly Detection

Rui Wang, Yeteng Wu, Xianling Zhang +1

Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learning approaches are fundamental…

cs.CV2026

Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving

Mengshi Qi, Xiaoyang Bi, Xianlin Zhang +1

Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather conditions due to sensor perception degradation. These challenges arise fr…

cs.CV2026

CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving

Zhaohong Liu, Hao Ye, Xianlin Zhang +1

End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised Fine-Tuning (SFT) often suffers from reasoning hallucinations an…

cs.CV2026

SGFormer++: Semantic Graph Transformer for Incremental 3D Scene Graph Generation

Mengshi Qi, Changsheng Lv, Zijian Fu +2

In this paper, we propose SGFormer++, a novel Semantic Graph Transformer for 3D scene graph generation (SGG), which aims to parse point cloud scenes into semantic structural graphs…

cs.CV2026

Global-Local Monte Carlo Tree Search in Vision-Language Models for Text-to-3D Indoor Scene Generation

Mengshi Qi, Wei Deng, Xianlin Zhang +1

Large Vision-Language Models have achieved significant reasoning performance in various tasks. However, there are few studies on text-to-3D indoor scene generation with LVLMs. The…

cs.CV2026

Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning

Mengshi Qi, Yeteng Wu, Wulian Yun +2

Evaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challenging in real-world scenarios. Howe…