4 papers
VLMine: Long-Tail Data Mining with Vision Language Models
Mao Ye, Gregory P. Meyer, Zaiwei Zhang +4
Ensuring robust performance on long-tail examples is an important problem for many real-world applications of machine learning, such as autonomous driving. This work focuses on the…
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
Zaiwei Zhang, Gregory P. Meyer, Zhichao Lu +3
For visual recognition, knowledge distillation typically involves transferring knowledge from a large, well-trained teacher model to a smaller student model. In this paper, we intr…
Cohere3D: Exploiting Temporal Coherence for Unsupervised Representation Learning of Vision-based Autonomous Driving
Yichen Xie, Hongge Chen, Gregory P. Meyer +6
Due to the lack of depth cues in images, multi-frame inputs are important for the success of vision-based perception, prediction, and planning in autonomous driving. Observations f…
SHIFT3D: Synthesizing Hard Inputs For Tricking 3D Detectors
Hongge Chen, Zhao Chen, Gregory P. Meyer +4
We present SHIFT3D, a differentiable pipeline for generating 3D shapes that are structurally plausible yet challenging to 3D object detectors. In safety-critical applications like…