6 papers · 1 filter
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
Qing Wang, Chong-Wah Ngo, Ee-Peng Lim
This paper addresses the challenges of learning representations for recipes and food images in the cross-modal retrieval problem. As the relationship between a recipe and its cooke…
LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
Qing Wang, Chong-Wah Ngo, Ee-Peng Lim +1
Training a model for food recognition is challenging because the training samples, which are typically crawled from the Internet, are visually different from the pictures captured…
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
Qing Wang, Chong-Wah Ngo, Yu Cao +1
Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food i…
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
Burak Satar, Zhixin Ma, Patrick A. Irawan +4
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultur…
Advancing Food Nutrition Estimation via Visual-Ingredient Feature Fusion
Huiyan Qi, Bin Zhu, Chong-Wah Ngo +2
Nutrition estimation is an important component of promoting healthy eating and mitigating diet-related health risks. Despite advances in tasks such as food classification and ingre…
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
Xiongwei Wu, Sicheng Yu, Ee-Peng Lim +1
In the realm of food computing, segmenting ingredients from images poses substantial challenges due to the large intra-class variance among the same ingredients, the emergence of n…