activity
20222024
most citedHyPoradise: An Open Baseline for Generative Speech Recognition with Large Language Models

7 citations · 31 across the 24 of their papers we have counts for

collaborators

37 papers

cs.SD2025

A correlation-permutation approach for speech-music encoders model merging

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jeremy H. M Wong +3

Creating a unified speech and music model requires expensive pre-training. Model merging can instead create an unified audio model with minimal computational expense. However, dire…

eess.AS2025

EASY: Emotion-aware Speaker Anonymization via Factorized Distillation

Jixun Yao, Hexin Liu, Eng Siong Chng +1

Emotion plays a significant role in speech interaction, conveyed through tone, pitch, and rhythm, enabling the expression of feelings and intentions beyond words to create a more p…

cs.CL2025

CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models

Sathya Krishnan Suresh, Tanmay Surana, Lim Zhi Hao +1

Code-switching (CS) poses a significant challenge for Large Language Models (LLMs), yet its comprehensibility remains underexplored in LLMs. We introduce CS-Sum, to evaluate the co…

cs.SD2025

Distilling a speech and music encoder with task arithmetic

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +4

Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding…

cs.SD2025

Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding

Dianwen Ng, Kun Zhou, Yi-Wen Chao +3

Achieving high-fidelity audio compression while preserving perceptual quality across diverse content remains a key challenge in Neural Audio Coding (NAC). We introduce MUFFIN, a fu…

cs.CL2025

Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages

Heqing Zou, Fengmao Lv, Desheng Zheng +2

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteri…