2 papers
cs.MM2024
RA-BLIP: Multimodal Adaptive Retrieval-Augmented Bootstrapping Language-Image Pre-training
Muhe Ding, Yang Ma, Pengda Qin +3
Multimodal Large Language Models (MLLMs) have recently received substantial interest, which shows their emerging potential as general-purpose models for various vision-language tas…
cs.CV2024
Preview-based Category Contrastive Learning for Knowledge Distillation
Muhe Ding, Jianlong Wu, Xue Dong +4
Knowledge distillation is a mainstream algorithm in model compression by transferring knowledge from the larger model (teacher) to the smaller model (student) to improve the perfor…