13 papers
ReGraph: Learning to Generate Recipe Graphs from Food Images
Guoshan Liu, Bin Zhu, Pengkun Jiao +3
Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.However, cooking is a structured transformation process in which in…
Enhancing Action and Ingredient Modeling for Semantically Grounded Recipe Generation
Guoshan Liu, Bin Zhu, Yian Li +3
Recent advances in Multimodal Large Language Models (MLMMs) have enabled recipe generation from food images, yet outputs often contain semantically incorrect actions or ingredients…
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
Qing Wang, Chong-Wah Ngo, Ee-Peng Lim
This paper addresses the challenges of learning representations for recipes and food images in the cross-modal retrieval problem. As the relationship between a recipe and its cooke…
LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
Qing Wang, Chong-Wah Ngo, Ee-Peng Lim +1
Training a model for food recognition is challenging because the training samples, which are typically crawled from the Internet, are visually different from the pictures captured…
Class Agnostic Instance-level Descriptor for Visual Instance Search
Qi-Ying Sun, Wan-Lei Zhao, Hui-Ying Xie +2
Despite the great success of the deep features in content-based image retrieval, the visual instance search remains challenging due to the lack of effective instance-level feature…
Efficient Test-Time Retrieval Augmented Generation
Hailong Yin, Bin Zhu, Jingjing Chen +1
Although Large Language Models (LLMs) demonstrate significant capabilities, their reliance on parametric knowledge often leads to inaccuracies. Retrieval Augmented Generation (RAG)…