2 papers
cs.CL2024
ASR Benchmarking: Need for a More Representative Conversational Dataset
Gaurav Maheshwari, Dmitry Ivanov, Théo Johannet +1
Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequatel…
cs.CL2024
Efficacy of Synthetic Data as a Benchmark
Gaurav Maheshwari, Dmitry Ivanov, Kevin El Haddad
Large language models (LLMs) have enabled a range of applications in zero-shot and few-shot learning settings, including the generation of synthetic datasets for training and testi…