7 papers
Learning to Scale Multilingual Representations for Vision-Language Tasks
Andrea Burns, Donghyun Kim, Derry Wijaya +2
Current multilingual vision-language models either require a large number of additional parameters for each supported language, or suffer performance degradation as languages are a…
Cross-domain Self-supervised Learning for Domain Adaptation with Few Source Labels
Donghyun Kim, Kuniaki Saito, Tae-Hyun Oh +3
Existing unsupervised domain adaptation methods aim to transfer knowledge from a label-rich source domain to an unlabeled target domain. However, obtaining labels for some source d…
LoGAN: Latent Graph Co-Attention Network for Weakly-Supervised Video Moment Retrieval
Reuben Tan, Huijuan Xu, Kate Saenko +1
The goal of weakly-supervised video moment retrieval is to localize the video segment most relevant to the given natural language query without access to temporal annotations durin…
MULE: Multimodal Universal Language Embedding
Donghyun Kim, Kuniaki Saito, Kate Saenko +2
Existing vision-language methods typically support two languages at a time at most. In this paper, we present a modular approach which can easily be incorporated into existing visi…
Learning Similarity Conditions Without Explicit Supervision
Reuben Tan, Mariya I. Vasileva, Kate Saenko +1
Many real-world tasks require models to compare images along multiple similarity conditions (e.g. similarity in color, category or shape). Existing methods often reason about these…
Language Features Matter: Effective Language Representations for Vision-Language Tasks
Andrea Burns, Reuben Tan, Kate Saenko +2
Shouldn't language and vision features be treated equally in vision-language (VL) tasks? Many VL approaches treat the language component as an afterthought, using simple language m…