1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024
Bringing Multimodality to Amazon Visual Search System
Xinliang Zhu, Michael Huang, Han Ding +10
Image to image matching has been well studied in the computer vision community. Previous studies mainly focus on training a deep metric learning model matching visual patterns betw…
cs.CV2024
DreamBlend: Advancing Personalized Fine-tuning of Text-to-Image Diffusion Models
Shwetha Ram, Tal Neiman, Qianli Feng +3
Given a small number of images of a subject, personalized image generation techniques can fine-tune large pre-trained text-to-image diffusion models to generate images of the subje…
cs.CV2024★ 1 cited
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
Sirnam Swetha, Jinyu Yang, Tal Neiman +5
Recent advancements in Multimodal Large Language Models (MLLMs) have revolutionized the field of vision-language understanding by integrating visual perception capabilities into La…