4 papers
Domain Generalizable Adaptation of 3D Vision-Language Models via Regularized Fine-Tuning
Sneha Paul, Zachary Patterson, Nizar Bouguila
Domain adaptation remains a central challenge in 3D vision, especially for multimodal foundation models that align 3D point clouds with visual and textual data. While these models…
An Adapter-free Fine-tuning Approach for Tuning 3D Foundation Models
Sneha Paul, Zachary Patterson, Nizar Bouguila
Point cloud foundation models demonstrate strong generalization, yet adapting them to downstream tasks remains challenging in low-data regimes. Full fine-tuning often leads to over…
Point Cloud as a Foreign Language for Multi-modal Large Language Model
Sneha Paul, Zachary Patterson, Nizar Bouguila
Multi-modal large language models (MLLMs) have shown remarkable progress in integrating visual and linguistic understanding. Recent efforts have extended these capabilities to 3D u…
LuKAN: A Kolmogorov-Arnold Network Framework for 3D Human Motion Prediction
Md Zahidul Hasan, A. Ben Hamza, Nizar Bouguila
The goal of 3D human motion prediction is to forecast future 3D poses of the human body based on historical motion data. Existing methods often face limitations in achieving a bala…