4.5k citations · 4.6k across the 16 of their papers we have counts for
1 paper · 2 filters
Yijun Shen, Delong Chen, Fan Liu +4
While densely annotated image captions significantly facilitate the learning of robust vision-language alignment, methodologies for systematically optimizing human annotation effor…