2 citations · 2 across the 4 of their papers we have counts for
4 papers
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
Georgios Paraskevopoulos, Chara Tsoukala, Athanasios Katsamanis +1
The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue i…
Weakly-supervised Automated Audio Captioning via text only training
Theodoros Kouzelis, Vassilis Katsouros
In recent years, datasets of paired audio and captions have enabled remarkable success in automatically generating descriptions for audio clips, namely Automated Audio Captioning (…
Investigating Personalization Methods in Text to Music Generation
Manos Plitsis, Theodoros Kouzelis, Georgios Paraskevopoulos +2
In this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the fir…
Weakly-supervised forced alignment of disfluent speech using phoneme-level modeling
Theodoros Kouzelis, Georgios Paraskevopoulos, Athanasios Katsamanis +1
The study of speech disorders can benefit greatly from time-aligned data. However, audio-text mismatches in disfluent speech cause rapid performance degradation for modern speech a…