1 paper
Xiwei Xuan, Xiaoqi Wang, Wenbin He +4
The advances in multi-modal foundation models (FMs) (e.g., CLIP and LLaVA) have facilitated the auto-labeling of large-scale datasets, enhancing model performance in challenging do…