Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Efficient ASR Training with Conversations that Never Happened
Máté Gedeon, Péter Mihajlik
Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that…
cs.CL2026
Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus
Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik +1
Conversational automatic speech recognition in Hungarian is constrained by the limited amount of publicly available dialogue-style training data. The BEA-Dialogue corpus addresses…
cs.CL2025
Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik +3
The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepr…