collaborators

10 papers

eess.AS2026

On the Role of Conversational Timing in Synthetic Training Data for ASR

Máté Gedeon, Péter Mihajlik

Synthetic multi-speaker conversations are widely used to train conversational automatic speech recognition (ASR) systems, but it remains unclear which timing properties make simula…

eess.AS2026

LibriConvo: Simulating Conversations from Read Literature for ASR and Diarization

Máté Gedeon, Péter Mihajlik

We introduce LibriConvo, a synthetic conversational speech corpus for speaker diarization and automatic speech recognition (ASR), built by instantiating the previously proposed Spe…

cs.CL2026

Efficient ASR Training with Conversations that Never Happened

Máté Gedeon, Péter Mihajlik

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that…

cs.CL2026

Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik +1

Conversational automatic speech recognition in Hungarian is constrained by the limited amount of publicly available dialogue-style training data. The BEA-Dialogue corpus addresses…

cs.SD2026

Speaker-Aware Simulation Improves Conversational Speech Recognition

Máté Gedeon, Péter Mihajlik

Automatic speech recognition (ASR) for conversational speech remains challenging due to the limited availability of large-scale, well-annotated multi-speaker dialogue data and the…

cs.CL2026

Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets

Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik +3

The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepr…