5 papers
Unsupervised Speech Recognition at the Syllable Level
Liming Wang, Kai-Wei Chang, Kunio Kashino +3
Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in…
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
Videet Mehta, Liming Wang, Hilde Kuehne +3
Large audio-language models (LALMs) exhibit strong zero-shot capabilities in multiple downstream tasks, such as audio question answering (AQA) and abstract reasoning; however, thes…
Towards Unsupervised Speech Recognition at the Syllable-Level
Liming Wang, Junrui Ni, Kai-Wei Chang +4
Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in…
Can Diffusion Models Disentangle? A Theoretical Perspective
Liming Wang, Muhammad Jehanzeb Mirza, Yishu Gong +6
This paper presents a novel theoretical framework for understanding how diffusion models can learn disentangled representations. Within this framework, we establish identifiability…
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
Edson Araujo, Andrew Rouditchenko, Yuan Gong +7
Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations t…