3 papers
cs.SD2025
A correlation-permutation approach for speech-music encoders model merging
Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jeremy H. M Wong +3
Creating a unified speech and music model requires expensive pre-training. Model merging can instead create an unified audio model with minimal computational expense. However, dire…
eess.AS2025
Multi-Distillation from Speech and Music Representation Models
Jui-Chiang Wei, Yi-Cheng Lin, Fabian Ritter-Gutierrez +1
Real-world audio often mixes speech and music, yet models typically handle only one domain. This paper introduces a multi-teacher distillation framework that unifies speech and mus…
cs.SD2025
Distilling a speech and music encoder with task arithmetic
Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +4
Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding…