4 papers
Reasoning-Driven Multimodal LLM for Domain Generalization
Zhipeng Xu, Zilong Wang, Xinyang Jiang +3
This paper addresses the domain generalization (DG) problem in deep learning. While most DG methods focus on enforcing visual feature invariance, we leverage the reasoning capabili…
DermINO: Hybrid Pretraining for a Versatile Dermatology Foundation Model
Jingkai Xu, De Cheng, Xiangqian Zhao +27
Skin diseases impose a substantial burden on global healthcare systems, driven by their high prevalence (affecting up to 70% of the population), complex diagnostic processes, and a…
Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
Yubin Wang, Xinyang Jiang, De Cheng +4
Visual prompt tuning offers significant advantages for adapting pre-trained visual foundation models to specific tasks. However, current research provides limited insight into the…
HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling
Yubin Wang, Xinyang Jiang, De Cheng +3
Prompt learning has become a prevalent strategy for adapting vision-language foundation models (VLMs) such as CLIP to downstream tasks. With the emergence of large language models…