activity
20242026
collaborators

9 papers

cs.SD2026

RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation

Rong Chao, Sung-Feng Huang, Moreno La Quatra +4

We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key…

cs.SD2026

One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications

Szu-Wei Fu, Rong Chao, Xuesong Yang +4

Different real-time speech applications impose distinct latency budgets, often requiring separately trained enhancement models for each scenario. In this paper, we propose a one-fo…

cs.SD2026

Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement

Szu-Wei Fu, Rong Chao, Xuesong Yang +6

Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in…

eess.AS2026

Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework

Hsiang-Cheng Yang, You-Jin Li, Rong Chao +3

This paper presents a Gaze-Guided Audio-Visual Speech Enhancement (GG-AVSE) framework to address the cocktail party problem. A major challenge in conventional AVSE is identifying t…

cs.SD2025

Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement

Rong Chao, Wenze Ren, You-Jin Li +5

Recent Mamba-based models have shown promise in speech enhancement by efficiently modeling long-range temporal dependencies. However, models like Speech Enhancement Mamba (SEMamba)…

cs.SD2025

Universal Speech Enhancement with Regression and Generative Mamba

Rong Chao, Rauf Nasretdinov, Yu-Chiang Frank Wang +3

The Interspeech 2025 URGENT Challenge aimed to advance universal, robust, and generalizable speech enhancement by unifying speech enhancement tasks across a wide variety of conditi…