activity
20192025
most citedAVTENet: A Human-Cognition-Inspired Audio-Visual Transformer-Based Ensemble Network for Video Deepfake Detection

19 citations · 93 across the 56 of their papers we have counts for

collaborators
Showing 2022 · eess.ASShow all

7 papers · 2 filters

eess.AS2022★ 1 cited

Mandarin Singing Voice Synthesis with Denoising Diffusion Probabilistic Wasserstein GAN

Yin-Ping Cho, Yu Tsao, Hsin-Min Wang +1

Singing voice synthesis (SVS) is the computer production of a human-like singing voice from given musical scores. To accomplish end-to-end SVS effectively and efficiently, this wor…

eess.AS2022

NASTAR: Noise Adaptive Speech Enhancement with Target-Conditional Resampling

Chi-Chang Lee, Cheng-Hung Hu, Yu-Chen Lin +3

For deep learning-based speech enhancement (SE) systems, the training-test acoustic mismatch can cause notable performance degradation. To address the mismatch issue, numerous nois…

eess.AS2022★ 3 cited

A Study of Using Cepstrogram for Countermeasure Against Replay Attacks

Shih-Kuang Lee, Yu Tsao, Hsin-Min Wang

This study investigated the cepstrogram properties and demonstrated their effectiveness as powerful countermeasures against replay attacks. A cepstrum analysis of replay attacks su…

eess.AS2022★ 2 cited

MTI-Net: A Multi-Target Speech Intelligibility Prediction Model

Ryandhimas E. Zezario, Szu-wei Fu, Fei Chen +3

Recently, deep learning (DL)-based non-intrusive speech assessment models have attracted great attention. Many studies report that these DL-based models yield satisfactory assessme…

eess.AS2022★ 1 cited

MBI-Net: A Non-Intrusive Multi-Branched Speech Intelligibility Prediction Model for Hearing Aids

Ryandhimas E. Zezario, Fei Chen, Chiou-Shann Fuh +2

Improving the user's hearing ability to understand speech in noisy environments is critical to the development of hearing aid (HA) devices. For this, it is important to derive a me…

eess.AS2022

Partially Fake Audio Detection by Self-attention-based Fake Span Discovery

Haibin Wu, Heng-Cheng Kuo, Naijun Zheng +5

The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly…