6 papers
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
Qing Wang, Chong-Wah Ngo, Ee-Peng Lim
This paper addresses the challenges of learning representations for recipes and food images in the cross-modal retrieval problem. As the relationship between a recipe and its cooke…
LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
Qing Wang, Chong-Wah Ngo, Ee-Peng Lim +1
Training a model for food recognition is challenging because the training samples, which are typically crawled from the Internet, are visually different from the pictures captured…
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
Qing Wang, Chong-Wah Ngo, Yu Cao +1
Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food i…
Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models
Bin Zhu, Yinxuan Gui, Huiyan Qi +3
Multimodal Large Language Models (MLLMs) have exhibited remarkable advancements in integrating different modalities, excelling in complex understanding and generation tasks. Despit…
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
Burak Satar, Zhixin Ma, Patrick A. Irawan +4
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultur…
Advancing Food Nutrition Estimation via Visual-Ingredient Feature Fusion
Huiyan Qi, Bin Zhu, Chong-Wah Ngo +2
Nutrition estimation is an important component of promoting healthy eating and mitigating diet-related health risks. Despite advances in tasks such as food classification and ingre…