26 citations · 36 across the 6 of their papers we have counts for
4 papers · 1 filter
FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
Yuehao Yin, Huiyan Qi, Bin Zhu +3
Large Multi-modal Models (LMMs) have made impressive progress in many vision-language tasks. Nevertheless, the performance of general LMMs in specific domains is still far from sat…
CgT-GAN: CLIP-guided Text GAN for Image Captioning
Jiarui Yu, Haoran Li, Yanbin Hao +3
The large-scale visual-language pre-trained model, Contrastive Language-Image Pre-training (CLIP), has significantly improved image captioning for scenarios without human-annotated…
Cross-lingual Adaptation for Recipe Retrieval with Mixup
Bin Zhu, Chong-Wah Ngo, Jingjing Chen +1
Cross-modal recipe retrieval has attracted research attention in recent years, thanks to the availability of large-scale paired data for training. Nevertheless, obtaining adequate…
Pyramid Fusion Dark Channel Prior for Single Image Dehazing
Qiyuan Liang, Bin Zhu, Chong-Wah Ngo
In this paper, we propose the pyramid fusion dark channel prior (PF-DCP) for single image dehazing. Based on the well-known Dark Channel Prior (DCP), we introduce an easy yet effec…