1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Ximing Wen, Mallika Mainali, Anik Sen
Vision Language Models (VLMs) have demonstrated strong reasoning capabilities in Visual Question Answering (VQA) tasks; however, their ability to perform Theory of Mind (ToM) tasks…