most citedTowards dialect-inclusive recognition in a low-resource language: are balanced corpora the answer?

2 citations · 2 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL2023

Zero-shot Audio Topic Reranking using Large Language Models

Mengjie Qian, Rao Ma, Adian Liusie +3

Multimodal Video Search by Examples (MVSE) investigates using video clips as the query term for information retrieval, rather than the more traditional text query. This enables far…

cs.CL2023

Towards spoken dialect identification of Irish

Liam Lonergan, Mengjie Qian, Neasa Ní Chiaráin +2

The Irish language is rich in its diversity of dialects and accents. This compounds the difficulty of creating a speech recognition system for the low-resource language, as such a…

cs.CL20232 cited

Towards dialect-inclusive recognition in a low-resource language: are balanced corpora the answer?

Liam Lonergan, Mengjie Qian, Neasa Ní Chiaráin +2

ASR systems are generally built for the spoken 'standard', and their performance declines for non-standard dialects/varieties. This is a problem for a language like Irish, where th…

cs.CL2023

Adapting an ASR Foundation Model for Spoken Language Assessment

Rao Ma, Mengjie Qian, Mark J. F. Gales +1

A crucial part of an accurate and reliable spoken language assessment system is the underlying ASR model. Recently, large-scale pre-trained ASR foundation models such as Whisper ha…

cs.CL2023

Can Generative Large Language Models Perform ASR Error Correction?

Rao Ma, Mengjie Qian, Potsawee Manakul +2

ASR error correction is an interesting option for post processing speech recognition system outputs. These error correction models are usually trained in a supervised fashion using…