activity
20162024
most citedAugmented Reality Meets Computer Vision : Efficient Data Generation for Urban Driving Scenes

16 citations · 23 across the 8 of their papers we have counts for

collaborators

11 papers

cs.CV2024★ 1 cited

VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision

Yi Xu, Yuxin Hu, Zaiwei Zhang +5

Human drivers rely on commonsense reasoning to navigate diverse and dynamic real-world scenarios. Existing end-to-end (E2E) autonomous driving (AD) models are typically optimized t…

cs.CV2024

VLMine: Long-Tail Data Mining with Vision Language Models

Mao Ye, Gregory P. Meyer, Zaiwei Zhang +4

Ensuring robust performance on long-tail examples is an important problem for many real-world applications of machine learning, such as autonomous driving. This work focuses on the…

cs.CV2023

ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts

Mu Cai, Haotian Liu, Dennis Park +4

While existing large vision-language multimodal models focus on whole image understanding, there is a prominent gap in achieving region-specific comprehension. Current approaches t…

cs.CV2023

NOVA: NOvel View Augmentation for Neural Composition of Dynamic Objects

Dakshit Agrawal, Jiajie Xu, Siva Karthik Mustikovela +3

We propose a novel-view augmentation (NOVA) strategy to train NeRFs for photo-realistic 3D composition of dynamic objects in a static scene. Compared to prior work, our framework s…

cs.CV2021

Self-Supervised Object Detection via Generative Image Synthesis

Siva Karthik Mustikovela, Shalini De Mello, Aayush Prakash +5

We present SSOD, the first end-to-end analysis-by synthesis framework with controllable GANs for the task of self-supervised object detection. We use collections of real world imag…

cs.CV2020★ 1 cited

Intrinsic Autoencoders for Joint Neural Rendering and Intrinsic Image Decomposition

Hassan Abu Alhaija, Siva Karthik Mustikovela, Justus Thies +4

Neural rendering techniques promise efficient photo-realistic image synthesis while at the same time providing rich control over scene parameters by learning the physical image for…