2 papers
cs.SD2026
Probing neural audio codecs for distinctions among English nuclear tunes
Juan Pablo Vigneaux, Jennifer Cole
State-of-the-art spoken dialogue models (Défossez et al. 2024; Schalkwyk et al. 2025) use neural audio codecs to "tokenize" audio signals into a lower-frequency stream of vectorial…
eess.AS2023
Crowdsourced and Automatic Speech Prominence Estimation
Max Morrison, Pranav Pawar, Nathan Pruyne +2
The prominence of a spoken word is the degree to which an average native listener perceives the word as salient or emphasized relative to its context. Speech prominence estimation…