16 citations · 23 across the 8 of their papers we have counts for
11 papers
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
Yi Xu, Yuxin Hu, Zaiwei Zhang +5
Human drivers rely on commonsense reasoning to navigate diverse and dynamic real-world scenarios. Existing end-to-end (E2E) autonomous driving (AD) models are typically optimized t…
VLMine: Long-Tail Data Mining with Vision Language Models
Mao Ye, Gregory P. Meyer, Zaiwei Zhang +4
Ensuring robust performance on long-tail examples is an important problem for many real-world applications of machine learning, such as autonomous driving. This work focuses on the…
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
Mu Cai, Haotian Liu, Dennis Park +4
While existing large vision-language multimodal models focus on whole image understanding, there is a prominent gap in achieving region-specific comprehension. Current approaches t…
NOVA: NOvel View Augmentation for Neural Composition of Dynamic Objects
Dakshit Agrawal, Jiajie Xu, Siva Karthik Mustikovela +3
We propose a novel-view augmentation (NOVA) strategy to train NeRFs for photo-realistic 3D composition of dynamic objects in a static scene. Compared to prior work, our framework s…
Self-Supervised Object Detection via Generative Image Synthesis
Siva Karthik Mustikovela, Shalini De Mello, Aayush Prakash +5
We present SSOD, the first end-to-end analysis-by synthesis framework with controllable GANs for the task of self-supervised object detection. We use collections of real world imag…
Intrinsic Autoencoders for Joint Neural Rendering and Intrinsic Image Decomposition
Hassan Abu Alhaija, Siva Karthik Mustikovela, Justus Thies +4
Neural rendering techniques promise efficient photo-realistic image synthesis while at the same time providing rich control over scene parameters by learning the physical image for…