1 paper
Suryaansh Jain, Rahasya Barkur, Vishal G +8
An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level captions, yet routinely miss t…