1 paper · 1 filter
Yusen Peng, Sachin Kumar
Recently, the advances in vision-language models, including contrastive pretraining and instruction tuning, have greatly pushed the frontier of multimodal AI. However, owing to the…