2 citations · 2 across the 2 of their papers we have counts for
3 papers
cs.CV2021
Multimodal Contrastive Training for Visual Representation Learning
Xin Yuan, Zhe Lin, Jason Kuen +5
We develop an approach to learning visual representations that embraces multimodal data, driven by a combination of intra- and inter-modal similarity preservation objectives. Unlik…
cs.CV2021
ALADIN: All Layer Adaptive Instance Normalization for Fine-grained Style Similarity
Dan Ruta, Saeid Motiian, Baldo Faieta +5
We present ALADIN (All Layer AdaIN); a novel architecture for searching images based on the similarity of their artistic style. Representation learning is critical to visual search…
cs.CV2019★ 2 cited
Multitask Text-to-Visual Embedding with Titles and Clickthrough Data
Pranav Aggarwal, Zhe Lin, Baldo Faieta +1
Text-visual (or called semantic-visual) embedding is a central problem in vision-language research. It typically involves mapping of an image and a text description to a common fea…