1 citations · 1 across the 1 of their papers we have counts for
3 papers
Test-Time Consistency in Vision Language Models
Shih-Han Chou, Shivam Chandhok, James J. Little +1
Vision-Language Models (VLMs) have achieved impressive performance across a wide range of multimodal tasks, yet they often exhibit inconsistent behavior when faced with semanticall…
MM-R: On (In-)Consistency of Vision-Language Models (VLMs)
Shih-Han Chou, Shivam Chandhok, James J. Little +1
With the advent of LLMs and variants, a flurry of research has emerged, analyzing the performance of such models across an array of tasks. While most studies focus on evaluating th…
Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)
Shih-Han Chou, Matthew Kowal, Yasmin Niknam +8
While progress has been made in the domain of video-language understanding, current state-of-the-art algorithms are still limited in their ability to understand videos at high leve…