119 citations · 183 across the 8 of their papers we have counts for
16 papers
i-Code: An Integrative and Composable Multimodal Learning Framework
Ziyi Yang, Yuwei Fang, Chenguang Zhu +17
Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to…
One model to enhance them all: array geometry agnostic multi-channel personalized speech enhancement
Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka +3
With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect…
Personalized Speech Enhancement: New Models and Comprehensive Evaluation
Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang +3
Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and…
UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data
Chengyi Wang, Yu Wu, Yao Qian +5
In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both unlabeled and labeled data, in which supervised phonetic CTC le…
Fusing Context Into Knowledge Graph for Commonsense Question Answering
Yichong Xu, Chenguang Zhu, Ruochen Xu +3
Commonsense question answering (QA) requires a model to grasp commonsense and factual knowledge to answer questions about world events. Many prior methods couple language modeling…
Mixed-Lingual Pre-training for Cross-lingual Summarization
Ruochen Xu, Chenguang Zhu, Yu Shi +2
Cross-lingual Summarization (CLS) aims at producing a summary in the target language for an article in the source language. Traditional solutions employ a two-step approach, i.e. t…