1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.SD2024
X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion
Chang Sun, Bo Qin
Target speaker extraction (TSE) is a technique for isolating a target speaker's voice from mixed speech using auxiliary features associated with the target speaker. It is another a…
cs.CV2024★ 1 cited
JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition
Chang Sun, Hong Yang, Bo Qin
Visual Speech Recognition (VSR) tasks are generally recognized to have a lower theoretical performance ceiling than Automatic Speech Recognition (ASR), owing to the inherent limita…