1 paper
Yijun Shen, Delong Chen, Fan Liu +4
While densely annotated image captions significantly facilitate the learning of robust vision-language alignment, methodologies for systematically optimizing human annotation effor…