4 citations · 4 across the 5 of their papers we have counts for
18 papers
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
Hyung Gun Chi, Zakaria Aldeneh, Tatiana Likhomanenko +5
We introduce DiceHuBERT, a knowledge distillation framework for compressing HuBERT, a widely used self-supervised learning (SSL)-based speech foundation model. Unlike existing dist…
A Variational Framework for Improving Naturalness in Generative Spoken Language Models
Li-Wei Chen, Takuya Higuchi, Zakaria Aldeneh +2
The success of large language models in text processing has inspired their adaptation to speech modeling. However, since speech is continuous and complex, it is often discretized f…
Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis
Akshita Gupta, Tatiana Likhomanenko, Karren Dai Yang +3
Unified decoder-only transformers have shown promise for multimodal generation, yet the mechanisms by which they synchronize modalities with heterogeneous sampling rates remain und…
Learning Spatially-Aware Language and Audio Embeddings
Bhavika Devnani, Skyler Seto, Zakaria Aldeneh +5
Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like "the lion roar came…
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
Zakaria Aldeneh, Takuya Higuchi, Jee-weon Jung +6
Iterative self-training, or iterative pseudo-labeling (IPL) -- using an improved model from the current iteration to provide pseudo-labels for the next iteration -- has proven to b…
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
Li-Wei Chen, Takuya Higuchi, He Bai +6
Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use…