8 papers
Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework
Pan Liao, Feng Yang, Di Wu +3
Semantic Multi-Object Tracking (SMOT) is evolving from purely geometric localization toward comprehensive video understanding. However, existing paradigms predominantly rely on clo…
Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving
Ziang Guo, Feng Yang, Xuefeng Zhang +6
Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat…
StereoMV2D: A Sparse Temporal Stereo-Enhanced Framework for Robust Multi-View 3D Object Detection
Di Wu, Feng Yang, Wenhui Zhao +4
Multi-view 3D object detection is a fundamental task in autonomous driving perception, where achieving a balance between detection accuracy and computational efficiency remains cru…
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
Xingcheng Zhou, Xuyuan Han, Feng Yang +3
We present OpenDriveVLA, a Vision Language Action model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially g…
VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning
Zishan Xu, Yifu Guo, Yuquan Lu +2
Traditional video reasoning segmentation methods rely on supervised fine-tuning, which limits generalization to out-of-distribution scenarios and lacks explicit reasoning. To addre…
FastTrackTr:Towards Fast Multi-Object Tracking with Transformers
Pan Liao, Feng Yang, Di Wu +3
Transformer-based multi-object tracking (MOT) methods have captured the attention of many researchers in recent years. However, these models often suffer from slow inference speeds…