paper

Edge Phoneme Recognition for Children's Speech through Age-Aware Training

arXiv:2608.10206

Abstract

Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a lightweight model to predict the age of the learner, as well as the phoneme sequence, enabled a 94M-parameter model to outperform WavLM Large models (317M) on the target DrivenData distribution, and fall within approximately 0.04 CER of competition ensembles with 90 times the parameters. This has enabled the creation of PhonemeTrainer, an application that can run on most modern cellular phones. This will ultimately enable better Automated Speech Recognition (ASR) and pronunciation helper apps for children's speech, with the privacy and compliance benefits that come with edge processing.

3 pages, 2 figures, 1 table. Demonstration paper presented at the non-archival demonstrations track of the 13th ACM Conference on Learning @ Scale (L@S '26), Seoul, South Korea, June 29-July 3, 2026. Not published in the ACM Digital Library

Edge Phoneme Recognition for Children's Speech through Age-Aware Training · wovepaper