3 citations · 3 across the 2 of their papers we have counts for
1 paper · 1 filter
Udith Haputhanthri, Declan Campbell, Rim Assouel +2
Despite success on standard benchmarks, vision language models display persistent failures on tasks involving processing of multi-object scenes, including many tasks that are relat…