collaborators

7 papers

cs.MM2026

Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis

Akshita Gupta, Tatiana Likhomanenko, Karren Dai Yang +3

Unified decoder-only transformers have shown promise for multimodal generation, yet the mechanisms by which they synchronize modalities with heterogeneous sampling rates remain und…

cs.LG2025

DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective

Hyung Gun Chi, Zakaria Aldeneh, Tatiana Likhomanenko +5

We introduce DiceHuBERT, a knowledge distillation framework for compressing HuBERT, a widely used self-supervised learning (SSL)-based speech foundation model. Unlike existing dist…

cs.CL2025

A Variational Framework for Improving Naturalness in Generative Spoken Language Models

Li-Wei Chen, Takuya Higuchi, Zakaria Aldeneh +2

The success of large language models in text processing has inspired their adaptation to speech modeling. However, since speech is continuous and complex, it is often discretized f…

cs.CL2025

dMel: Speech Tokenization made Simple

Richard He Bai, Tatiana Likhomanenko, Ruixiang Zhang +3

Large language models have revolutionized natural language processing by leveraging self-supervised pretraining on vast textual data. Inspired by this success, researchers have inv…

eess.AS2025

Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels

Zakaria Aldeneh, Takuya Higuchi, Jee-weon Jung +6

Iterative self-training, or iterative pseudo-labeling (IPL) -- using an improved model from the current iteration to provide pseudo-labels for the next iteration -- has proven to b…

eess.AS2025

Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models

Li-Wei Chen, Takuya Higuchi, He Bai +6

Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use…