activity
20162023
most citedLearning Deep Representations of Medical Images using Siamese CNNs with Application to Content-Based Image Retrieval

62 citations · 319 across the 16 of their papers we have counts for

collaborators
Showing cs.CLShow all

16 papers · 1 filter

cs.CL2023

CoLLD: Contrastive Layer-to-layer Distillation for Compressing Multilingual Pre-trained Speech Encoders

Heng-Jui Chang, Ning Dong, Ruslan Mavlyutov +2

Large-scale self-supervised pre-trained speech encoders outperform conventional approaches in speech recognition and translation tasks. Due to the high cost of developing these lar…

cs.CL202314 cited

SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Seamless Communication, Loïc Barrault, Yu-An Chung +65

What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed…

cs.CL20225 cited

Speech-to-Speech Translation For A Real-world Unwritten Language

Peng-Jen Chen, Kevin Tran, Yilin Yang +13

We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard te…

cs.CL202150 cited

SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training

Ankur Bapna, Yu-an Chung, Nan Wu +7

Unsupervised pre-training is now the predominant approach for both text and speech understanding. Self-attention models pre-trained on large amounts of unannotated data have been h…

cs.CL2020

Non-Autoregressive Predictive Coding for Learning Speech Representations from Local Dependencies

Alexander H. Liu, Yu-An Chung, James Glass

Self-supervised speech representations have been shown to be effective in a variety of speech applications. However, existing representation learning methods generally rely on the…

cs.CL2020

SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding

Yu-An Chung, Chenguang Zhu, Michael Zeng

Spoken language understanding (SLU) requires a model to analyze input acoustic signal to understand its linguistic content and make predictions. To boost the models' performance, v…