2 papers
cs.CV2025
Rethinking Visual Information Processing in Multimodal LLMs
Dongwan Kim, Viresh Ranjan, Takashi Nagata +2
Despite the remarkable success of the LLaVA architecture for vision-language tasks, its design inherently struggles to effectively integrate visual features due to the inherent mis…
cs.CV2024
Bringing Multimodality to Amazon Visual Search System
Xinliang Zhu, Michael Huang, Han Ding +10
Image to image matching has been well studied in the computer vision community. Previous studies mainly focus on training a deep metric learning model matching visual patterns betw…