4 papers
ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +3
Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically pr…
Multi-Distillation from Speech and Music Representation Models
Jui-Chiang Wei, Yi-Cheng Lin, Fabian Ritter-Gutierrez +1
Real-world audio often mixes speech and music, yet models typically handle only one domain. This paper introduces a multi-teacher distillation framework that unifies speech and mus…
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang +77
Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spo…
Distilling a speech and music encoder with task arithmetic
Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +4
Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding…