2 papers
cs.CL2024
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
Yang Yuhang, Peng Yizhou, Eng Siong Chng +1
The integration of large language models (LLMs) with pre-trained speech models has opened up new avenues in automatic speech recognition (ASR). While LLMs excel in multimodal under…
eess.AS2024
Room Impulse Responses help attackers to evade Deep Fake Detection
Hieu-Thi Luong, Duc-Tuan Truong, Kong Aik Lee +1
The ASVspoof 2021 benchmark, a widely-used evaluation framework for anti-spoofing, consists of two subsets: Logical Access (LA) and Deepfake (DF), featuring samples with varied cod…