activity
20192022
most citedMetricGAN-U: Unsupervised speech enhancement/ dereverberation based only on noisy/ reverberated speech

8 citations · 14 across the 3 of their papers we have counts for

collaborators

8 papers

eess.AS2022

Partially Fake Audio Detection by Self-attention-based Fake Span Discovery

Haibin Wu, Heng-Cheng Kuo, Naijun Zheng +5

The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly…

cs.SD20218 cited

MetricGAN-U: Unsupervised speech enhancement/ dereverberation based only on noisy/ reverberated speech

Szu-Wei Fu, Cheng Yu, Kuo-Hsuan Hung +2

Most of the deep learning-based speech enhancement models are learned in a supervised manner, which implies that pairs of noisy and clean speech are required during training. Conse…

eess.AS2021

EMA2S: An End-to-End Multimodal Articulatory-to-Speech System

Yu-Wen Chen, Kuo-Hsuan Hung, Shang-Yi Chuang +4

Synthesized speech from articulatory movements can have real-world use for patients with vocal cord disorders, situations requiring silent speech, or in high-noise environments. In…

eess.SP2020

Deep Learning Based Signal Enhancement of Low-Resolution Accelerometer for Fall Detection Systems

Kai-Chun Liu, Kuo-Hsuan Hung, Chia-Yeh Hsieh +3

In the last two decades, fall detection (FD) systems have been developed as a popular assistive technology. Such systems automatically detect critical fall events and immediately a…

eess.AS2020

A Study of Incorporating Articulatory Movement Information in Speech Enhancement

Yu-Wen Chen, Kuo-Hsuan Hung, Shang-Yi Chuang +3

Although deep learning algorithms are widely used for improving speech enhancement (SE) performance, the performance remains limited under highly challenging conditions, such as un…

eess.AS20206 cited

Waveform-based Voice Activity Detection Exploiting Fully Convolutional networks with Multi-Branched Encoders

Cheng Yu, Kuo-Hsuan Hung, I-Fan Lin +3

In this study, we propose an encoder-decoder structured system with fully convolutional networks to implement voice activity detection (VAD) directly on the time-domain waveform. T…