151 citations · 587 across the 49 of their papers we have counts for
90 papers
Cascaded Multilingual Audio-Visual Learning from Videos
Andrew Rouditchenko, Angie Boggust, David Harwath +8
In this paper, we explore self-supervised audio-visual models that learn from instructional videos. Prior work has shown that these models can relate spoken words and sounds to vis…
On the Interplay Between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis
Cheng-I Jeff Lai, Erica Cooper, Yang Zhang +8
Are end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a sta…
Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset
Ian Palmer, Andrew Rouditchenko, Andrei Barbu +2
Visually-grounded spoken language datasets can enable models to learn cross-modal correspondences with very weak supervision. However, modern audio-visual datasets contain biases t…
An Empirical Study on Few-shot Knowledge Probing for Pretrained Language Models
Tianxing He, Kyunghyun Cho, James Glass
Prompt-based knowledge probing for 1-hop relations has been used to measure how much world knowledge is stored in pretrained language models. Existing work uses considerable amount…
Interpretable Propaganda Detection in News Articles
Seunghak Yu, Giovanni Da San Martino, Mitra Mohtarami +2
Online users today are exposed to misleading and propagandistic news articles and media posts on a daily basis. To counter thus, a number of approaches have been designed aiming to…
Mitigating Biases in Toxic Language Detection through Invariant Rationalization
Yung-Sung Chuang, Mingye Gao, Hongyin Luo +4
Automatic detection of toxic language plays an essential role in protecting social media users, especially minority groups, from verbal abuse. However, biases toward some attribute…