3 papers
cs.SD2024
Efficient Streaming LLM for Speech Recognition
Junteng Jia, Gil Keren, Wei Zhou +6
Recent works have shown that prompting large language models with audio encodings can unlock speech recognition capabilities. However, existing techniques do not scale efficiently,…
cs.SD2024
Utilizing Speaker Profiles for Impersonation Audio Detection
Hao Gu, JiangYan Yi, Chenglong Wang +5
Fake audio detection is an emerging active topic. A growing number of literatures have aimed to detect fake utterance, which are mostly generated by Text-to-speech (TTS) or voice c…
cs.SD2023
TST: Time-Sparse Transducer for Automatic Speech Recognition
Xiaohui Zhang, Mangui Liang, Zhengkun Tian +2
End-to-end model, especially Recurrent Neural Network Transducer (RNN-T), has achieved great success in speech recognition. However, transducer requires a great memory footprint an…