activity
20222026
most citedComputation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2026

PhiNet: Speaker Verification with Phonetic Interpretability

Yi Ma, Shuai Wang, Tianchi Liu +1

Despite remarkable progress, automatic speaker verification (ASV) systems typically lack the transparency required for high-accountability applications. Motivated by how human expe…

eess.AS2025

Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive Front-ends

Qiquan Zhang, Buddhi Wickramasinghe, Eliathamby Ambikairajah +2

Hand-crafted features, such as Mel-filterbanks, have traditionally been the choice for many audio processing applications. Recently, there has been a growing interest in learnable…

eess.AS2022

Self-Transriber: Few-shot Lyrics Transcription with Self-training

Xiaoxue Gao, Xianghu Yue, Haizhou Li

The current lyrics transcription approaches heavily rely on supervised learning with labeled data, but such data are scarce and manual labeling of singing is expensive. How to bene…

eess.AS2022

PoLyScriber: Integrated Fine-tuning of Extractor and Lyrics Transcriber for Polyphonic Music

Xiaoxue Gao, Chitralekha Gupta, Haizhou Li

Lyrics transcription of polyphonic music is challenging as the background music affects lyrics intelligibility. Typically, lyrics transcription can be performed by a two-step pipel…

eess.AS2022

Music-robust Automatic Lyrics Transcription of Polyphonic Music

Xiaoxue Gao, Chitralekha Gupta, Haizhou Li

Lyrics transcription of polyphonic music is challenging because singing vocals are corrupted by the background music. To improve the robustness of lyrics transcription to the backg…