2 papers
cs.CV2026
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
Zhenyu Ning, Guangda Liu, Qihao Jin +4
Recent developments in Video Large Language Models (Video LLMs) have enabled models to process hour-long videos and exhibit exceptional performance. Nonetheless, the Key-Value (KV)…
cs.RO2026
Unveiling the Surprising Efficacy of Navigation Understanding in End-to-End Autonomous Driving
Zhihua Hua, Junli Wang, Pengfei LI +6
Global navigation information and local scene understanding are two crucial components of autonomous driving systems. However, our experimental results indicate that many end-to-en…