1 paper · 1 filter
Jihoon Kwon, Kyle Min, Jy-yong Sohn
Despite recent advances, vision-language models trained with standard contrastive objectives still struggle with compositional reasoning -- the ability to understand structured rel…