4 citations · 9 across the 3 of their papers we have counts for
3 papers
cs.CV2022★ 2 cited
FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and Captioning
Suvir Mirchandani, Licheng Yu, Mengjiao Wang +4
Multimodal tasks in the fashion domain have significant potential for e-commerce, but involve challenging vision-and-language learning problems - e.g., retrieving a fashion item gi…
cs.CV2022★ 3 cited
CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni Retrieval
Licheng Yu, Jun Chen, Animesh Sinha +4
We introduce CommerceMM - a multimodal model capable of providing a diverse and granular understanding of commerce topics associated to the given piece of content (image, text, ima…
cs.CV2021★ 4 cited
Large-Scale Attribute-Object Compositions
Filip Radenovic, Animesh Sinha, Albert Gordo +2
We study the problem of learning how to predict attribute-object compositions from images, and its generalization to unseen compositions missing from the training data. To the best…