1 citations · 1 across the 4 of their papers we have counts for
4 papers
Spatial-Temporal Activity-Informed Diarization and Separation
Yicheng Hsu, Ssuhan Chen, Mingsian R. Bai
A robust multichannel speaker diarization and separation system is proposed by exploiting the spatio-temporal activity of the speakers. The system is realized in a hybrid architect…
Deep Beamforming for Speech Enhancement and Speaker Localization with an Array Response-Aware Loss Function
Hsinyu Chang, Yicheng Hsu, Mingsian R. Bai
Recent research advances in deep neural network (DNN)-based beamformers have shown great promise for speech enhancement under adverse acoustic conditions. Different network archite…
Array Configuration-Agnostic Personal Voice Activity Detection Based on Spatial Coherence
Yicheng Hsu, Mingsian R. Bai
Personal voice activity detection has received increased attention due to the growing popularity of personal mobile devices and smart speakers. PVAD is often an integral element to…
Multi-channel target speech enhancement based on ERB-scaled spatial coherence features
Yicheng Hsu, Yonghan Lee, Mingsian R. Bai
Recently, speech enhancement technologies that are based on deep learning have received considerable research attention. If the spatial information in microphone signals is exploit…