1 paper · 1 filter
Guoshan Liu, Bin Zhu, Yian Li +3
Recent advances in Multimodal Large Language Models (MLMMs) have enabled recipe generation from food images, yet outputs often contain semantically incorrect actions or ingredients…