7 citations · 29 across the 23 of their papers we have counts for
26 papers · 1 filter
Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
Haolin He, Xingjian Du, Renhe Sun +16
Large Audio Language Models (LALMs) represent an important frontier in multimodal AI, addressing diverse audio tasks. Recently, post-training of LALMs has received increasing atten…
An Investigation on Applying Acoustic Feature Conversion to ASR of Adult and Child Speech
Wei Liu, Jingyu Li, Tan Lee
The performance of child speech recognition is generally less satisfactory compared to adult speech due to limited amount of training data. Significant performance degradation is e…
Data Augmentation with Locally-time Reversed Speech for Automatic Speech Recognition
Si-Ioi Ng, Tan Lee
Psychoacoustic studies have shown that locally-time reversed (LTR) speech, i.e., signal samples time-reversed within a short segment, can be accurately recognised by human listener…
A study on the efficacy of model pre-training in developing neural text-to-speech system
Guangyan Zhang, Yichong Leng, Daxin Tan +5
In the development of neural text-to-speech systems, model pre-training with a large amount of non-target speakers' data is a common approach. However, in terms of ultimately achie…
Improving Text-Independent Speaker Verification with Auxiliary Speakers Using Graph
Jingyu Li, Si-Ioi Ng, Tan Lee
The paper presents a novel approach to refining similarity scores between input utterances for robust speaker verification. Given the embeddings from a pair of input utterances, a…
Utterance-level neural confidence measure for end-to-end children speech recognition
Wei Liu, Tan Lee
Confidence measure is a performance index of particular importance for automatic speech recognition (ASR) systems deployed in real-world scenarios. In the present study, utterance-…