activity
20242026
collaborators

6 papers

eess.AS2026

Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving

Ziang Guo, Feng Yang, Xuefeng Zhang +6

Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat…

cs.CV2025

StereoMV2D: A Sparse Temporal Stereo-Enhanced Framework for Robust Multi-View 3D Object Detection

Di Wu, Feng Yang, Wenhui Zhao +4

Multi-view 3D object detection is a fundamental task in autonomous driving perception, where achieving a balance between detection accuracy and computational efficiency remains cru…

cs.CV2025

VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning

Zishan Xu, Yifu Guo, Yuquan Lu +2

Traditional video reasoning segmentation methods rely on supervised fine-tuning, which limits generalization to out-of-distribution scenarios and lacks explicit reasoning. To addre…

cs.CV2025

OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

Xingcheng Zhou, Xuyuan Han, Feng Yang +3

We present OpenDriveVLA, a Vision Language Action model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially g…

cs.CV2024

HV-BEV: Decoupling Horizontal and Vertical Feature Sampling for Multi-View 3D Object Detection

Di Wu, Feng Yang, Benlian Xu +3

The application of vision-based multi-view environmental perception system has been increasingly recognized in autonomous driving technology, especially the BEV-based models. Curre…

cs.CV2024

FastTrackTr:Towards Fast Multi-Object Tracking with Transformers

Pan Liao, Feng Yang, Di Wu +3

Transformer-based multi-object tracking (MOT) methods have captured the attention of many researchers in recent years. However, these models often suffer from slow inference speeds…