15 citations · 35 across the 18 of their papers we have counts for
23 papers · 1 filter
Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?
Apratim Bhattacharyya, Shweta Mahajan, Sanjay Haresh +5
Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large lang…
Generative Scenario Rollouts for End-to-End Autonomous Driving
Rajeev Yasarla, Deepti Hegde, Shizhong Han +10
Vision-Language-Action (VLA) models are emerging as highly effective planning models for end-to-end autonomous driving systems. However, current works mostly rely on imitation lear…
DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos
Rajeev Yasarla, Shizhong Han, Hong Cai +1
Camera-based 3D object detection in Bird's Eye View (BEV) is one of the most important perception tasks in autonomous driving. Earlier methods rely on dense BEV features, which are…
RoCA: Robust Cross-Domain End-to-End Autonomous Driving
Rajeev Yasarla, Shizhong Han, Hsin-Pai Cheng +7
End-to-end (E2E) autonomous driving has recently emerged as a new paradigm, offering significant potential. However, few studies have looked into the practical challenge of deploym…
Distilling Multi-modal Large Language Models for Autonomous Driving
Deepti Hegde, Rajeev Yasarla, Hong Cai +7
Autonomous driving demands safe motion planning, especially in critical "long-tail" scenarios. Recent end-to-end autonomous driving systems leverage large language models (LLMs) as…
ToSA: Token Selective Attention for Efficient Vision Transformers
Manish Kumar Singh, Rajeev Yasarla, Hong Cai +2
In this paper, we propose a novel token selective attention approach, ToSA, which can identify tokens that need to be attended as well as those that can skip a transformer layer. M…