activity
20172022
most citedImproving One-Shot Learning through Fusing Side Information

43 citations · 133 across the 15 of their papers we have counts for

collaborators

28 papers

cs.LG20221 cited

Greedy Modality Selection via Approximate Submodular Maximization

Runxiang Cheng, Gargi Balasubramaniam, Yifei He +2

Multimodal learning considers learning from multi-modality data, aiming to fuse heterogeneous sources of information. However, it is not always feasible to leverage all available m…

cs.CV2022

Towards Multimodal Multitask Scene Understanding Models for Indoor Mobile Agents

Yao-Hung Hubert Tsai, Hanlin Goh, Ali Farhadi +1

The perception system in personalized mobile agents requires developing indoor scene understanding models, which can understand 3D geometries, capture objectiveness, analyze human…

cs.CV2022

Paraphrasing Is All You Need for Novel Object Captioning

Cheng-Fu Yang, Yao-Hung Hubert Tsai, Wan-Cyuan Fan +3

Novel object captioning (NOC) aims to describe images containing objects without observing their ground truth captions during training. Due to the absence of caption annotation, ca…

cs.LG20225 cited

Conditional Contrastive Learning with Kernel

Yao-Hung Hubert Tsai, Tianqin Li, Martin Q. Ma +4

Conditional contrastive learning frameworks consider the conditional sampling procedure that constructs positive or negative data pairs conditioned on specific variables. Fair cont…

cs.LG20222 cited

Learning Weakly-Supervised Contrastive Representations

Yao-Hung Hubert Tsai, Tianqin Li, Weixin Liu +3

We argue that a form of the valuable information provided by the auxiliary information is its implied data clustering information. For instance, considering hashtags as auxiliary i…

cs.CL202127 cited

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai +3

Self-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no l…