1 paper
Xiaomeng Jin, Jeonghwan Kim, Yu Zhou +4
In Multimodal Language Models (MLMs), the cost of manually annotating high-quality image-text pair data for fine-tuning and alignment is extremely high. While existing multimodal d…