collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Efficient ASR Training with Conversations that Never Happened

Máté Gedeon, Péter Mihajlik

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that…

cs.CL2026

Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik +1

Conversational automatic speech recognition in Hungarian is constrained by the limited amount of publicly available dialogue-style training data. The BEA-Dialogue corpus addresses…

cs.CL2026

Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets

Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik +3

The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepr…

cs.CL2025

A Comparative Analysis of Static Word Embeddings for Hungarian

Máté Gedeon

This paper presents a comprehensive analysis of various static word embeddings for Hungarian, including traditional models such as Word2Vec, FastText, as well as static embeddings…

cs.CL2025

Retrieval-Enhanced Few-Shot Prompting for Speech Event Extraction

Máté Gedeon

Speech Event Extraction (SpeechEE) is a challenging task that lies at the intersection of Automatic Speech Recognition (ASR) and Natural Language Processing (NLP), requiring the id…