2 citations · 2 across the 4 of their papers we have counts for
4 papers · 1 filter
Rethinking Visual Information Processing in Multimodal LLMs
Dongwan Kim, Viresh Ranjan, Takashi Nagata +2
Despite the remarkable success of the LLaVA architecture for vision-language tasks, its design inherently struggles to effectively integrate visual features due to the inherent mis…
Bringing Multimodality to Amazon Visual Search System
Xinliang Zhu, Michael Huang, Han Ding +10
Image to image matching has been well studied in the computer vision community. Previous studies mainly focus on training a deep metric learning model matching visual patterns betw…
Leveraging Large Language Models for Multimodal Search
Oriol Barbany, Michael Huang, Xinliang Zhu +1
Multimodal search has become increasingly important in providing users with a natural and effective way to ex-press their search intentions. Images offer fine-grained details of th…
ProcSim: Proxy-based Confidence for Robust Similarity Learning
Oriol Barbany, Xiaofan Lin, Muhammet Bastan +1
Deep Metric Learning (DML) methods aim at learning an embedding space in which distances are closely related to the inherent semantic similarity of the inputs. Previous studies hav…