6 papers
Visual-Informed Speech Enhancement Using Attention-Based Beamforming
Chihyun Liu, Jiaxuan Fan, Mingtung Sun +3
Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance.…
AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
Sahibzada Adil Shahzad, Ammarah Hashmi, Yan-Tsung Peng +2
Multimodal manipulations (also known as audio-visual deepfakes) make it difficult for unimodal deepfake detectors to detect forgeries in multimedia content. To avoid the spread of…
NGGAN: Noise Generation GAN Based on the Practical Measurement Dataset for Narrowband Powerline Communications
Ying-Ren Chien, Po-Heng Chou, You-Jie Peng +3
To effectively process impulse noise for narrowband powerline communications (NB-PLCs) transceivers, capturing comprehensive statistics of nonperiodic asynchronous impulsive noise…
MECKD: Deep Learning-Based Fall Detection in Multilayer Mobile Edge Computing With Knowledge Distillation
Wei-Lung Mao, Chun-Chi Wang, Po-Heng Chou +2
The rising aging population has increased the importance of fall detection (FD) systems as an assistive technology, where deep learning techniques are widely applied to enhance acc…
Deep Reinforcement Learning-Based Precoding for Multi-RIS-Aided Multiuser Downlink Systems with Practical Phase Shift
Po-Heng Chou, Bo-Ren Zheng, Wan-Jen Huang +3
This study considers multiple reconfigurable intelligent surfaces (RISs)-aided multiuser downlink systems with the goal of jointly optimizing the transmitter precoding and RIS phas…
From Evaluation to Optimization: Neural Speech Assessment for Downstream Applications
Yu Tsao
The evaluation of synthetic and processed speech has long been a cornerstone of audio engineering and speech science. Although subjective listening tests remain the gold standard f…