3 papers
cs.CL2026
Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus
Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik +1
Conversational automatic speech recognition in Hungarian is constrained by the limited amount of publicly available dialogue-style training data. The BEA-Dialogue corpus addresses…
cs.CL2025
Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik +3
The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepr…
eess.AS2022
BEA-Base: A Benchmark for ASR of Spontaneous Hungarian
P. Mihajlik, A. Balog, T. E. Gráczi +3
Hungarian is spoken by 15 million people, still, easily accessible Automatic Speech Recognition (ASR) benchmark datasets - especially for spontaneous speech - have been practically…