6 citations · 9 across the 4 of their papers we have counts for
5 papers
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
Jing-Xuan Zhang, Genshun Wan, Jianqing Gao +1
Audio-visual representation learning is crucial for advancing multimodal speech processing tasks, such as lipreading and audio-visual speech recognition. Recently, speech foundatio…
Is Lip Region-of-Interest Sufficient for Lipreading?
Jing-Xuan Zhang, Gen-Shun Wan, Jia Pan
Lip region-of-interest (ROI) is conventionally used for visual input in the lipreading task. Few works have adopted the entire face as visual input because lip-excluded parts of th…
Voice Conversion by Cascading Automatic Speech Recognition and Text-to-Speech Synthesis with Prosody Transfer
Jing-Xuan Zhang, Li-Juan Liu, Yan-Nian Chen +4
With the development of automatic speech recognition (ASR) and text-to-speech synthesis (TTS) technique, it's intuitive to construct a voice conversion system by cascading an ASR a…
Recognition-Synthesis Based Non-Parallel Voice Conversion with Adversarial Learning
Jing-Xuan Zhang, Zhen-Hua Ling, Li-Rong Dai
This paper presents an adversarial learning method for recognition-synthesis based non-parallel voice conversion. A recognizer is used to transform acoustic features into linguisti…
ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang, Junichi Yamagishi, Massimiliano Todisco +37
Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition. Unfortunately, just like all other biometric systems, ASV is v…