4 papers
LiquidTAD: Efficient Temporal Action Detection via Parallel Liquid-Inspired Temporal Relaxation
Zepeng Sun, Naichuan Zheng, Hailun Xia +3
Temporal Action Detection (TAD) requires precise localization of action boundaries within long, untrimmed video sequences. While current high-performing methods achieve strong accu…
SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving
Jingyu Li, Junjie Wu, Dongnan Hu +6
Recent end-to-end autonomous driving approaches have leveraged Vision-Language Models (VLMs) to enhance planning capabilities in complex driving scenarios. However, VLMs are inhere…
FLARE: Learning Future-Aware Latent Representations from Vision-Language Models for Autonomous Driving
Chengen Xie, Chonghao Sima, Tianyu Li +4
While Vision-Language Models (VLMs) offer rich world knowledge for end-to-end autonomous driving, current approaches heavily rely on labor-intensive language annotations (e.g., VQA…
GaussianAD: Gaussian-Centric End-to-End Autonomous Driving
Wenzhao Zheng, Junjie Wu, Yao Zheng +8
Vision-based autonomous driving shows great potential due to its satisfactory performance and low costs. Most existing methods adopt dense representations (e.g., bird's eye view) o…