25 citations · 33 across the 18 of their papers we have counts for
19 papers
Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?
Apratim Bhattacharyya, Shweta Mahajan, Sanjay Haresh +5
Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large lang…
Generative Scenario Rollouts for End-to-End Autonomous Driving
Rajeev Yasarla, Deepti Hegde, Shizhong Han +10
Vision-Language-Action (VLA) models are emerging as highly effective planning models for end-to-end autonomous driving systems. However, current works mostly rely on imitation lear…
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
Shweta Mahajan, Shreya Kadambi, Hoang Le +4
We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transformations driven by real-world a…
Attention Guided Alignment in Efficient Vision-Language Models
Shweta Mahajan, Hoang Le, Hyojin Park +3
Large Vision-Language Models (VLMs) rely on effective multimodal alignment between pre-trained vision encoders and Large Language Models (LLMs) to integrate visual and textual info…
RoCA: Robust Cross-Domain End-to-End Autonomous Driving
Rajeev Yasarla, Shizhong Han, Hsin-Pai Cheng +7
End-to-end (E2E) autonomous driving has recently emerged as a new paradigm, offering significant potential. However, few studies have looked into the practical challenge of deploym…
DuoLoRA : Cycle-consistent and Rank-disentangled Content-Style Personalization
Aniket Roy, Shubhankar Borse, Shreya Kadambi +8
We tackle the challenge of jointly personalizing content and style from a few examples. A promising approach is to train separate Low-Rank Adapters (LoRA) and merge them effectivel…