27 citations · 93 across the 15 of their papers we have counts for
8 papers · 1 filter
DMCL: Distillation Multiple Choice Learning for Multimodal Action Recognition
Nuno C. Garcia, Sarah Adel Bargal, Vitaly Ablavsky +3
In this work, we address the problem of learning an ensemble of specialist networks using multimodal data, while considering the realistic and challenging scenario of possible miss…
Siamese Natural Language Tracker: Tracking by Natural Language Descriptions with Siamese Trackers
Qi Feng, Vitaly Ablavsky, Qinxun Bai +1
We propose a novel Siamese Natural Language Tracker (SNLT), which brings the advancements in visual tracking to the tracking by natural language (NL) descriptions task. The propose…
MULE: Multimodal Universal Language Embedding
Donghyun Kim, Kuniaki Saito, Kate Saenko +2
Existing vision-language methods typically support two languages at a time at most. In this paper, we present a modular approach which can easily be incorporated into existing visi…
Language Features Matter: Effective Language Representations for Vision-Language Tasks
Andrea Burns, Reuben Tan, Kate Saenko +2
Shouldn't language and vision features be treated equally in vision-language (VL) tasks? Many VL approaches treat the language component as an afterthought, using simple language m…
Real-time Visual Object Tracking with Natural Language Description
Qi Feng, Vitaly Ablavsky, Qinxun Bai +2
In recent years, deep-learning-based visual object trackers have been studied thoroughly, but handling occlusions and/or rapid motion of the target remains challenging. In this wor…
Weakly-supervised Compositional FeatureAggregation for Few-shot Recognition
Ping Hu, Ximeng Sun, Kate Saenko +1
Learning from a few examples is a challenging task for machine learning. While recent progress has been made for this problem, most of the existing methods ignore the compositional…