4 papers
Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
Qixin Deng, Bryan Pardo, Thrasyvoulos N Pappas
Understanding and modeling the relationship between language and sound are essential for applications such as music information retrieval, text-guided music generation, and audio c…
Fine-Grained and Interpretable Neural Speech Editing
Max Morrison, Cameron Churchwell, Nathan Pruyne +1
Fine-grained editing of speech attributes$\unicode{x2014}$such as prosody (i.e., the pitch, loudness, and phoneme durations), pronunciation, speaker identity, and formants$\unicode…
High-Fidelity Neural Phonetic Posteriorgrams
Cameron Churchwell, Max Morrison, Bryan Pardo
A phonetic posteriorgram (PPG) is a time-varying categorical distribution over acoustic units of speech (e.g., phonemes). PPGs are a popular representation in speech generation due…
Crowdsourced and Automatic Speech Prominence Estimation
Max Morrison, Pranav Pawar, Nathan Pruyne +2
The prominence of a spoken word is the degree to which an average native listener perceives the word as salient or emphasized relative to its context. Speech prominence estimation…