collaborators

8 papers

cs.CV2026

CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios

Sparsh Garg, Yi-Wen Chen, Vijay Kumar B G +1

While modern autonomous driving systems excel at perception tasks such as object detection and trajectory prediction, they lack the high-level causal reasoning required to interpre…

cs.CV2026

Driving Video Retrieval for Complex Queries with Structured Grounding

Manyi Yao, Sparsh Garg, Christian Shelton +2

Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins a…

cs.CV2026

What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs

Abhishek Aich, Sparsh Garg, Vijay Kumar BG +2

Driving vision-language models (VLMs) must accurately understand scenes across diverse conditions defined by Operational Design Domains (ODDs), yet verification remains sparse: man…

cs.CV2026

HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes

Mauricio Soroco, Francesco Pittaluga, Zaid Tasneem +5

Ensuring safety in autonomous driving requires scalable generation of realistic, controllable driving scenes beyond what real-world testing provides. Yet existing instruction guide…

cs.CV2026

Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation

Manyi Yao, Abhishek Aich, Yumin Suh +3

Vision transformer based models bring significant improvements for image segmentation tasks. Although these architectures offer powerful capabilities irrespective of specific segme…

cs.CV2025

iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning

Manyi Yao, Bingbing Zhuang, Sparsh Garg +4

Grounding large language models (LLMs) in domain-specific tasks like post-hoc dash-cam driving video analysis is challenging due to their general-purpose training and lack of struc…