3 papers
cs.CL2025
Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models
Bin Zhu, Yinxuan Gui, Huiyan Qi +3
Multimodal Large Language Models (MLLMs) have exhibited remarkable advancements in integrating different modalities, excelling in complex understanding and generation tasks. Despit…
cs.CL2025
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan +48
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate th…
cs.MM2025
Efficient Prompt Tuning for Hierarchical Ingredient Recognition
Yinxuan Gui, Bin Zhu, Jingjing Chen +1
Fine-grained ingredient recognition presents a significant challenge due to the diverse appearances of ingredients, resulting from different cutting and cooking methods. While exis…