13 citations · 15 across the 7 of their papers we have counts for
3 papers · 1 filter
NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
Chaofan Tao, Gukyeong Kwon, Varad Gunjal +7
We study the capability of Video-Language (VidL) models in understanding compositions between objects, attributes, actions and their relations. Composition understanding becomes pa…
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
Prannay Kaul, Zhizhong Li, Hao Yang +4
Mitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which…
Learning Expressive Prompting With Residuals for Vision Transformers
Rajshekhar Das, Yonatan Dukler, Avinash Ravichandran +1
Prompt learning is an efficient approach to adapt transformers by inserting learnable set of parameters into the input and intermediate representations of a pre-trained model. In t…