speech processing

An Empirical Recipe for Universal Phone Recognition

arXiv:2603.29042

summary

The paper introduces PhoneticXEUS, a multilingual phone recognition system trained on large-scale data that achieves state-of-the-art error rates on both many languages and accented English, and provides an empirical recipe detailing the effects of data scale, SSL representations, and training objectives.

Abstract

Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive. Highly performant English-focused models do not generalize across languages, while multilingual models underutilize pretrained representations. It also remains unclear how data scale, architecture, and training objective contribute to multilingual PR. We present PhoneticXEUS -- trained on large-scale multilingual data and achieving state-of-the-art performance on both multilingual (17.7% PFER) and accented English speech (10.6% PFER). Through controlled ablations with evaluations across 100+ languages under a unified scheme, we empirically establish our training recipe and quantify the impact of SSL representations, data scale, and loss objectives. In addition, we analyze error patterns across language families, accented speech, and articulatory features. All data and code are released openly at https://github.com/changelinglab/PhoneticXeus

Accepted at Interspeech 2026. Code: https://github.com/changelinglab/PhoneticXeus

Topics & keywords

#phone recognition#multilingual speech#self-supervised learning#low-resource languages#acoustic modelingPhoneticXEUSself-supervised learninglarge-scale multilingual dataphone error rateSSL representationstraining loss objectives
An Empirical Recipe for Universal Phone Recognition · wovepaper