80 citations · 216 across the 14 of their papers we have counts for
16 papers · 1 filter
End-to-End Visual Editing with a Generatively Pre-Trained Artist
Andrew Brown, Cheng-Yang Fu, Omkar Parkhi +2
We consider the targeted image editing problem: blending a region in a source image with a driver image that specifies the desired change. Differently from prior works, we solve th…
LoopITR: Combining Dual and Cross Encoder Architectures for Image-Text Retrieval
Jie Lei, Xinlei Chen, Ning Zhang +4
Dual encoders and cross encoders have been widely used for image-text retrieval. Between the two, the dual encoder encodes the image and text independently followed by a dot produc…
CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni Retrieval
Licheng Yu, Jun Chen, Animesh Sinha +4
We introduce CommerceMM - a multimodal model capable of providing a diverse and granular understanding of commerce topics associated to the given piece of content (image, text, ima…
VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
Linjie Li, Jie Lei, Zhe Gan +12
Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily…
Large-Scale Attribute-Object Compositions
Filip Radenovic, Animesh Sinha, Albert Gordo +2
We study the problem of learning how to predict attribute-object compositions from images, and its generalization to unseen compositions missing from the training data. To the best…
Connecting What to Say With Where to Look by Modeling Human Attention Traces
Zihang Meng, Licheng Yu, Ning Zhang +4
We introduce a unified framework to jointly model images, text, and human attention traces. Our work is built on top of the recent Localized Narratives annotation framework [30], w…