2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2025
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
Monika Wysoczańska, Shyamal Buch, Anurag Arnab +1
Large vision-language models (VLMs) often struggle to generate long and factual captions. However, traditional measures for hallucination and factuality are not well suited for eva…
cs.CV2024★ 2 cited
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
Gagan Jain, Nidhi Hegde, Aditya Kusupati +5
The visual medium (images and videos) naturally contains a large amount of information redundancy, thereby providing a great opportunity for leveraging efficiency in processing. Wh…
cs.CV2024
Streaming Dense Video Captioning
Xingyi Zhou, Anurag Arnab, Shyamal Buch +5
An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos, predict rich, detailed textual descr…