5 papers · 1 filter
Efficient ASR Training with Conversations that Never Happened
Máté Gedeon, Péter Mihajlik
Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that…
Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus
Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik +1
Conversational automatic speech recognition in Hungarian is constrained by the limited amount of publicly available dialogue-style training data. The BEA-Dialogue corpus addresses…
Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik +3
The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepr…
A Comparative Analysis of Static Word Embeddings for Hungarian
Máté Gedeon
This paper presents a comprehensive analysis of various static word embeddings for Hungarian, including traditional models such as Word2Vec, FastText, as well as static embeddings…
Retrieval-Enhanced Few-Shot Prompting for Speech Event Extraction
Máté Gedeon
Speech Event Extraction (SpeechEE) is a challenging task that lies at the intersection of Automatic Speech Recognition (ASR) and Natural Language Processing (NLP), requiring the id…