18 citations · 18 across the 2 of their papers we have counts for
3 papers
cs.SD2022
The Volcspeech system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge
Chen Shen, Yi Liu, Wenzhi Fan +6
This paper describes our submission to ICASSP 2022 Multi-channel Multi-party Meeting Transcription (M2MeT) Challenge. For Track 1, we propose several approaches to empower the clus…
cs.LG2020★ 18 cited
Speech-Image Semantic Alignment Does Not Depend on Any Prior Classification Tasks
Masood S. Mortazavi
Semantically-aligned datasets can be used to explore "visually-grounded speech". In a majority of existing investigations, features of an image signal are extract…
eess.SP2018
Efficient improvement of frequency-domain Kalman filter
Wenzhi Fan, Kai Chen, Jing Lu +1
The frequency-domain Kalman filter (FKF) has been utilized in many audio signal processing applications due to its fast convergence speed and robustness. However, the performance o…