collaborators

8 papers

cs.CL2026

A framework for analyzing concept representations in neural models

Burin Naowarat, Hao Tang, Sharon Goldwater

Understanding how neural models represent human-interpretable concepts is challenging. Prior work has explored linear concept subspaces from diverse perspectives, such as probing a…

cs.CL2025

Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models

Michele Gubian, Ioana Krehan, Oli Liu +2

Analyses of self-supervised speech models have begun to reveal where and how they represent different types of information. However, almost all analyses have focused on English. He…

cs.SD2025

Effective Context in Neural Speech Models

Yen Meng, Sharon Goldwater, Hao Tang

Modern neural speech models benefit from having longer context, and many approaches have been proposed to increase the maximum context a model can use. However, few have attempted…

cs.CL2025

Revisiting Common Assumptions about Arabic Dialects in NLP

Amr Keleg, Sharon Goldwater, Walid Magdy

Arabic has diverse dialects, where one dialect can be substantially different from the others. In the NLP literature, some assumptions about these dialects are widely adopted (e.g.…

cs.CL2024

Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations

Mukhtar Mohamed, Oli Danyi Liu, Hao Tang +1

Self-supervised speech representations can hugely benefit downstream speech technologies, yet the properties that make them useful are still poorly understood. Two candidate proper…

cs.CL2024

Estimating the Level of Dialectness Predicts Interannotator Agreement in Multi-dialect Arabic Datasets

Amr Keleg, Walid Magdy, Sharon Goldwater

On annotating multi-dialect Arabic datasets, it is common to randomly assign the samples across a pool of native Arabic speakers. Recent analyses recommended routing dialectal samp…