2 papers
cs.CV2024
You Only Speak Once to See
Wenhao Yang, Jianguo Wei, Wenhuan Lu +1
Grounding objects in images using visual cues is a well-established approach in computer vision, yet the potential of audio as a modality for object recognition and grounding remai…
eess.AS2024
Integrated Multi-Level Knowledge Distillation for Enhanced Speaker Verification
Wenhao Yang, Jianguo Wei, Wenhuan Lu +2
Knowledge distillation (KD) is widely used in audio tasks, such as speaker verification (SV), by transferring knowledge from a well-trained large model (the teacher) to a smaller,…