5 papers · 1 filter
CLAP: Unsupervised 3D Representation Learning for Fusion 3D Perception via Curvature Sampling and Prototype Learning
Runjian Chen, Hang Zhang, Avinash Ravichandran +4
Unsupervised 3D representation learning reduces the burden of labeling multimodal 3D data for fusion perception tasks. Among different pre-training paradigms, differentiable-render…
Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation
Reza Akbarian Bafghi, Carden Bagwell, Avinash Ravichandran +2
Adapting deep learning models to new domains often requires computationally intensive retraining and risks catastrophic forgetting. While fine-tuning enables domain-specific adapta…
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
Zaiwei Zhang, Gregory P. Meyer, Zhichao Lu +3
For visual recognition, knowledge distillation typically involves transferring knowledge from a large, well-trained teacher model to a smaller student model. In this paper, we intr…
InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models
Nirat Saini, Navaneeth Bodla, Ashish Shrivastava +4
We introduce InVi, an approach for inserting or replacing objects within videos (referred to as inpainting) using off-the-shelf, text-to-image latent diffusion models. InVi targets…
GenMM: Geometrically and Temporally Consistent Multimodal Data Generation for Video and LiDAR
Bharat Singh, Viveka Kulharia, Luyu Yang +3
Multimodal synthetic data generation is crucial in domains such as autonomous driving, robotics, augmented/virtual reality, and retail. We propose a novel approach, GenMM, for join…