collaborators

6 papers

cs.CV2026

Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval

Qing Wang, Chong-Wah Ngo, Ee-Peng Lim

This paper addresses the challenges of learning representations for recipes and food images in the cross-modal retrieval problem. As the relationship between a recipe and its cooke…

cs.CV2025

LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets

Qing Wang, Chong-Wah Ngo, Ee-Peng Lim +1

Training a model for food recognition is challenging because the training samples, which are typically crawled from the Internet, are visually different from the pictures captured…

cs.CV2025

Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval

Qing Wang, Chong-Wah Ngo, Yu Cao +1

Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food i…

cs.CL2025

Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models

Bin Zhu, Yinxuan Gui, Huiyan Qi +3

Multimodal Large Language Models (MLLMs) have exhibited remarkable advancements in integrating different modalities, excelling in complex understanding and generation tasks. Despit…

cs.CV2025

Seeing Culture: A Benchmark for Visual Reasoning and Grounding

Burak Satar, Zhixin Ma, Patrick A. Irawan +4

Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultur…

cs.CV2025

Advancing Food Nutrition Estimation via Visual-Ingredient Feature Fusion

Huiyan Qi, Bin Zhu, Chong-Wah Ngo +2

Nutrition estimation is an important component of promoting healthy eating and mitigating diet-related health risks. Despite advances in tasks such as food classification and ingre…