Showing cs.SDShow all
3 papers · 1 filter
cs.SD2025
Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token
Shogo Seki, Shaoxiang Dang, Li Li
The Dilated FAVOR Conformer (DF-Conformer) is an efficient variant of the Conformer architecture designed for speech enhancement (SE). It employs fast attention through positive or…
cs.SD2024
U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation
Shaoxiang Dang, Tetsuya Matsumoto, Yoshinori Takeuchi +1
The topic of speech separation involves separating mixed speech with multiple overlapping speakers into several streams, with each stream containing speech from only one speaker. M…
cs.SD2024
Developing vocal system impaired patient-aimed voice quality assessment approach using ASR representation-included multiple features
Shaoxiang Dang, Tetsuya Matsumoto, Yoshinori Takeuchi +7
The potential of deep learning in clinical speech processing is immense, yet the hurdles of limited and imbalanced clinical data samples loom large. This article addresses these ch…