activity
20232025
collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2025

ASTAR-NTU solution to AudioMOS Challenge 2025 Track1

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +3

Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically pr…

cs.SD2025

A correlation-permutation approach for speech-music encoders model merging

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jeremy H. M Wong +3

Creating a unified speech and music model requires expensive pre-training. Model merging can instead create an unified audio model with minimal computational expense. However, dire…

cs.SD2025

Distilling a speech and music encoder with task arithmetic

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +4

Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding…

cs.SD2024

Dataset-Distillation Generative Model for Speech Emotion Recognition

Fabian Ritter-Gutierrez, Kuan-Po Huang, Jeremy H. M Wong +4

Deep learning models for speech rely on large datasets, presenting computational challenges. Yet, performance hinges on training data size. Dataset Distillation (DD) aims to learn…

cs.SD2023

Noise robust distillation of self-supervised speech models via correlation metrics

Fabian Ritter-Gutierrez, Kuan-Po Huang, Dianwen Ng +4

Compared to large speech foundation models, small distilled models exhibit degraded noise robustness. The student's robustness can be improved by introducing noise at the inputs du…