activity
20172025
most citedSpatial LibriSpeech: An Augmented Dataset for Spatial Audio Learning

4 citations · 4 across the 5 of their papers we have counts for

collaborators

18 papers

cs.LG2025

DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective

Hyung Gun Chi, Zakaria Aldeneh, Tatiana Likhomanenko +5

We introduce DiceHuBERT, a knowledge distillation framework for compressing HuBERT, a widely used self-supervised learning (SSL)-based speech foundation model. Unlike existing dist…

cs.CL2025

A Variational Framework for Improving Naturalness in Generative Spoken Language Models

Li-Wei Chen, Takuya Higuchi, Zakaria Aldeneh +2

The success of large language models in text processing has inspired their adaptation to speech modeling. However, since speech is continuous and complex, it is often discretized f…

cs.MM2024

Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis

Akshita Gupta, Tatiana Likhomanenko, Karren Dai Yang +3

Unified decoder-only transformers have shown promise for multimodal generation, yet the mechanisms by which they synchronize modalities with heterogeneous sampling rates remain und…

cs.SD2024

Learning Spatially-Aware Language and Audio Embeddings

Bhavika Devnani, Skyler Seto, Zakaria Aldeneh +5

Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like "the lion roar came…

eess.AS2024

Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels

Zakaria Aldeneh, Takuya Higuchi, Jee-weon Jung +6

Iterative self-training, or iterative pseudo-labeling (IPL) -- using an improved model from the current iteration to provide pseudo-labels for the next iteration -- has proven to b…

eess.AS2024

Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models

Li-Wei Chen, Takuya Higuchi, He Bai +6

Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use…