9 papers
VTO: Visual Tool Orchestration for Video Anomaly Detection
Rui Wang, Yeteng Wu, Xianling Zhang +1
Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learning approaches are fundamental…
Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving
Mengshi Qi, Xiaoyang Bi, Xianlin Zhang +1
Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather conditions due to sensor perception degradation. These challenges arise fr…
CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving
Zhaohong Liu, Hao Ye, Xianlin Zhang +1
End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised Fine-Tuning (SFT) often suffers from reasoning hallucinations an…
SGFormer++: Semantic Graph Transformer for Incremental 3D Scene Graph Generation
Mengshi Qi, Changsheng Lv, Zijian Fu +2
In this paper, we propose SGFormer++, a novel Semantic Graph Transformer for 3D scene graph generation (SGG), which aims to parse point cloud scenes into semantic structural graphs…
Global-Local Monte Carlo Tree Search in Vision-Language Models for Text-to-3D Indoor Scene Generation
Mengshi Qi, Wei Deng, Xianlin Zhang +1
Large Vision-Language Models have achieved significant reasoning performance in various tasks. However, there are few studies on text-to-3D indoor scene generation with LVLMs. The…
Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
Mengshi Qi, Yeteng Wu, Wulian Yun +2
Evaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challenging in real-world scenarios. Howe…