collaborators

6 papers

eess.AS2026

FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching

Fabian Ritter-Gutierrez, Md Asif Jalal, Pablo Peso Parada +5

Whispered-to-normal (W2N) speech conversion aims to reconstruct missing phonation from whispered input while preserving content and speaker identity. This task is challenging due t…

cs.SD2025

ASTAR-NTU solution to AudioMOS Challenge 2025 Track1

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +3

Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically pr…

cs.SD2025

A correlation-permutation approach for speech-music encoders model merging

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jeremy H. M Wong +3

Creating a unified speech and music model requires expensive pre-training. Model merging can instead create an unified audio model with minimal computational expense. However, dire…

eess.AS2025

Multi-Distillation from Speech and Music Representation Models

Jui-Chiang Wei, Yi-Cheng Lin, Fabian Ritter-Gutierrez +1

Real-world audio often mixes speech and music, yet models typically handle only one domain. This paper introduces a multi-teacher distillation framework that unifies speech and mus…

cs.CL2025

Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang +77

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spo…

cs.SD2025

Distilling a speech and music encoder with task arithmetic

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +4

Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding…