activity
20242026
collaborators

8 papers

cs.CL2026

"This Wasn't Made for Me": Recentering User Experience and Emotional Impact in the Evaluation of ASR Bias

Siyu Liang, Alicia Beckford Wassink

Studies on bias in Automatic Speech Recognition (ASR) tend to focus on reporting error rates for speakers of underrepresented dialects, yet less research examines the human side of…

cs.CL2025

A Sociophonetic Analysis of Racial Bias in Commercial ASR Systems Using the Pacific Northwest English Corpus

Michael Scott, Siyu Liang, Alicia Wassink +1

This paper presents a systematic evaluation of racial bias in four major commercial automatic speech recognition (ASR) systems using the Pacific Northwest English (PNWE) corpus. We…

cs.CL2025

The Limits of Data Scaling: Sub-token Utilization and Acoustic Saturation in Multilingual ASR

Siyu Liang, Nicolas Ballier, Gina-Anne Levow +1

How much audio is needed to fully observe a multilingual ASR model's learned sub-token inventory across languages, and does data disparity in multilingual pre-training affect how t…

cs.CL2025

The Tonogenesis Continuum in Tibetan: A Computational Investigation

Siyu Liang, Zhaxi Zerong

Tonogenesis-the historical process by which segmental contrasts evolve into lexical tone-has traditionally been studied through comparative reconstruction and acoustic phonetics. W…

cs.CL2025

Beyond WER: Probing Whisper's Sub-token Decoder Across Diverse Language Resource Levels

Siyu Liang, Nicolas Ballier, Gina-Anne Levow +1

While large multilingual automatic speech recognition (ASR) models achieve remarkable performance, the internal mechanisms of the end-to-end pipeline, particularly concerning fairn…

cs.CL2025

Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages

Siyu Liang, Gina-Anne Levow

Automatic Speech Recognition (ASR) has reached impressive accuracy for high-resource languages, yet its utility in linguistic fieldwork remains limited. Recordings collected in fie…