4 papers
P3T: Prototypical Point-level Prompt Tuning with Enhanced Generalization for 3D Vision-Language Models
Geunyoung Jung, Soohong Kim, Kyungwoo Song +1
With the rise of pre-trained models in the 3D point cloud domain for a wide range of real-world applications, adapting them to downstream tasks has become increasingly important. H…
Generating Accurate and Detailed Captions for High-Resolution Images
Hankyeol Lee, Gawon Seo, Kyounggyu Lee +3
Vision-language models (VLMs) often struggle to generate accurate and detailed captions for high-resolution images since they are typically pre-trained on low-resolution inputs (e.…
Enhancing Visual Classification using Comparative Descriptors
Hankyeol Lee, Gawon Seo, Wonseok Choi +3
The performance of vision-language models (VLMs), such as CLIP, in visual classification tasks, has been enhanced by leveraging semantic knowledge from large language models (LLMs)…
DaWin: Training-free Dynamic Weight Interpolation for Robust Adaptation
Changdae Oh, Yixuan Li, Kyungwoo Song +2
Adapting a pre-trained foundation model on downstream tasks should ensure robustness against distribution shifts without the need to retrain the whole model. Although existing weig…