4 papers
M-QUEST -- Meme Question-Understanding Evaluation on Semantics and Toxicity
Stefano De Giorgis, Ting-Chih Chen, Filip Ilievski
Internet memes are a powerful form of online communication, yet their nature and reliance on commonsense knowledge make toxicity detection challenging. Identifying key features for…
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara +1
Multimodal Large Language Models (MLLMs) have experienced rapid progress in visual recognition tasks in recent years. Given their potential integration into many critical applicati…
COLUMBUS: Evaluating COgnitive Lateral Understanding through Multiple-choice reBUSes
Koen Kraaijveld, Yifan Jiang, Kaixin Ma +1
While visual question-answering (VQA) benchmarks have catalyzed the development of reasoning techniques, they have focused on vertical thinking. Effective problem-solving also nece…
Robust Text Classification: Analyzing Prototype-Based Networks
Zhivar Sourati, Darshan Deshpande, Filip Ilievski +2
Downstream applications often require text classification models to be accurate and robust. While the accuracy of the state-of-the-art Language Models (LMs) approximates human perf…