Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
PROFIT: A Specialized Optimizer for Deep Fine Tuning
Anirudh S Chakravarthy, Shuai Kyle Zheng, Xin Huang +4
The fine-tuning of pre-trained models has become ubiquitous in generative AI, computer vision, and robotics. Although much attention has been paid to improving the efficiency of fi…
cs.CV2024
VLMine: Long-Tail Data Mining with Vision Language Models
Mao Ye, Gregory P. Meyer, Zaiwei Zhang +4
Ensuring robust performance on long-tail examples is an important problem for many real-world applications of machine learning, such as autonomous driving. This work focuses on the…
cs.CV2024
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
Mu Cai, Haotian Liu, Dennis Park +4
While existing large vision-language multimodal models focus on whole image understanding, there is a prominent gap in achieving region-specific comprehension. Current approaches t…