3 papers
cs.CV2026
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
Zhenyu Ning, Guangda Liu, Qihao Jin +4
Recent developments in Video Large Language Models (Video LLMs) have enabled models to process hour-long videos and exhibit exceptional performance. Nonetheless, the Key-Value (KV)…
cs.RO2026
Unveiling the Surprising Efficacy of Navigation Understanding in End-to-End Autonomous Driving
Zhihua Hua, Junli Wang, Pengfei LI +6
Global navigation information and local scene understanding are two crucial components of autonomous driving systems. However, our experimental results indicate that many end-to-en…
cs.RO2024
Dual-AEB: Synergizing Rule-Based and Multimodal Large Language Models for Effective Emergency Braking
Wei Zhang, Pengfei Li, Junli Wang +8
Automatic Emergency Braking (AEB) systems are a crucial component in ensuring the safety of passengers in autonomous vehicles. Conventional AEB systems primarily rely on closed-set…