activity
20152022
most citedVisual Madlibs: Fill in the blank Image Generation and Question Answering

80 citations · 216 across the 14 of their papers we have counts for

collaborators

21 papers

cs.CV20221 cited

End-to-End Visual Editing with a Generatively Pre-Trained Artist

Andrew Brown, Cheng-Yang Fu, Omkar Parkhi +2

We consider the targeted image editing problem: blending a region in a source image with a driver image that specifies the desired change. Differently from prior works, we solve th…

cs.CV20226 cited

LoopITR: Combining Dual and Cross Encoder Architectures for Image-Text Retrieval

Jie Lei, Xinlei Chen, Ning Zhang +4

Dual encoders and cross encoders have been widely used for image-text retrieval. Between the two, the dual encoder encodes the image and text independently followed by a dot produc…

cs.CV20223 cited

CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni Retrieval

Licheng Yu, Jun Chen, Animesh Sinha +4

We introduce CommerceMM - a multimodal model capable of providing a diverse and granular understanding of commerce topics associated to the given piece of content (image, text, ima…

cs.CL2021

MTVR: Multilingual Moment Retrieval in Videos

Jie Lei, Tamara L. Berg, Mohit Bansal

We introduce mTVR, a large-scale multilingual video moment retrieval dataset, containing 218K English and Chinese queries from 21.8K TV show video clips. The dataset is collected b…

cs.CV202138 cited

VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

Linjie Li, Jie Lei, Zhe Gan +12

Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily…

cs.CV20214 cited

Large-Scale Attribute-Object Compositions

Filip Radenovic, Animesh Sinha, Albert Gordo +2

We study the problem of learning how to predict attribute-object compositions from images, and its generalization to unseen compositions missing from the training data. To the best…