13 papers
Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis
Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang +4
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over p…
HAMMER: Harmonic-Aware Parallel Context Modeling and Discriminator-Free Perceptual Optimization for Speech Enhancement
Shang-Fu Chen, Szu-Wei Fu, Sung-Feng Huang +3
Recent speech enhancement systems combine self-attention and Mamba to capture global interactions and long-range dependencies. Yet these hybrids usually operate as sequence mixers…
RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation
Rong Chao, Sung-Feng Huang, Moreno La Quatra +4
We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key…
One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications
Szu-Wei Fu, Rong Chao, Xuesong Yang +4
Different real-time speech applications impose distinct latency budgets, often requiring separately trained enhancement models for each scenario. In this paper, we propose a one-fo…
Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement
Szu-Wei Fu, Rong Chao, Xuesong Yang +6
Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in…
How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
Ke-Han Lu, Szu-Wei Fu, Chao-Han Huck Yang +13
Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-…