109 citations · 256 across the 29 of their papers we have counts for
13 papers · 1 filter
Audio-Visual Event Recognition through the lens of Adversary
Juncheng B Li, Kaixin Ma, Shuhui Qu +2
As audio/visual classification models are widely deployed for sensitive tasks like content filtering at scale, it is critical to understand their robustness along with improving th…
Multimodal Speech Recognition with Unstructured Audio Masking
Tejas Srinivasan, Ramon Sanabria, Florian Metze +1
Visual context has been shown to be useful for automatic speech recognition (ASR) systems when the speech signal is noisy or corrupted. Previous work, however, has only demonstrate…
On Long-Tailed Phenomena in Neural Machine Translation
Vikas Raunak, Siddharth Dalmia, Vivek Gupta +1
State-of-the-art Neural Machine Translation (NMT) models struggle with generating low-frequency tokens, tackling which remains a major challenge. The analysis of long-tailed phenom…
Fine-Grained Grounding for Multimodal Speech Recognition
Tejas Srinivasan, Ramon Sanabria, Florian Metze +1
Multimodal automatic speech recognition systems integrate information from images to improve speech recognition quality, by grounding the speech in the visual context. While visual…
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano +4
The dominant paradigm for learning video-text representations -- noise contrastive learning -- increases the similarity of the representations of pairs of samples that are known to…
Revisiting Factorizing Aggregated Posterior in Learning Disentangled Representations
Ze Cheng, Juncheng Li, Chenxu Wang +4
In the problem of learning disentangled representations, one of the promising methods is to factorize aggregated posterior by penalizing the total correlation of sampled latent var…