papers

Publications (8)

cs.CV2021

MLIM: Vision-and-Language Model Pre-training with Masked Language and Image Modeling

Tarik Arici, Mehmet Saygin Seyfioglu, Tal Neiman +5

Vision-and-Language Pre-training (VLP) improves model performance for downstream tasks that require image and text inputs. Current VLP approaches differ on (i) model architecture (…

cs.CV2024

Bringing Multimodality to Amazon Visual Search System

Xinliang Zhu, Michael Huang, Han Ding +10

Image to image matching has been well studied in the computer vision community. Previous studies mainly focus on training a deep metric learning model matching visual patterns betw…

cs.LG2021

Hypergraph Pre-training with Graph Neural Networks

Boxin Du, Changhe Yuan, Robert Barton +2

Despite the prevalence of hypergraphs in a variety of high-impact applications, there are relatively few works on hypergraph representation learning, most of which primarily focus…

cs.LG2026

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs

Jinqi Luo, Jinyu Yang, Tal Neiman +5

Multimodal Large Language Models (MLLMs) have been shown to be vulnerable to malicious queries that can elicit unsafe responses. Recent work uses prompt engineering, response class…

cs.CV2024

DreamBlend: Advancing Personalized Fine-tuning of Text-to-Image Diffusion Models

Shwetha Ram, Tal Neiman, Qianli Feng +3

Given a small number of images of a subject, personalized image generation techniques can fine-tune large pre-trained text-to-image diffusion models to generate images of the subje…

cs.LG2021

Graph Neural Networks for Inconsistent Cluster Detection in Incremental Entity Resolution

Robert A. Barton, Tal Neiman, Changhe Yuan

Online stores often utilize product relationships such as bundles and substitutes to improve their catalog quality and guide customers through myriad choices. Entity resolution usi…