2 citations · 2 across the 3 of their papers we have counts for
3 papers
Weakly-supervised Automated Audio Captioning via text only training
Theodoros Kouzelis, Vassilis Katsouros
In recent years, datasets of paired audio and captions have enabled remarkable success in automatically generating descriptions for audio clips, namely Automated Audio Captioning (…
Investigating Personalization Methods in Text to Music Generation
Manos Plitsis, Theodoros Kouzelis, Georgios Paraskevopoulos +2
In this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the fir…
Weakly-supervised forced alignment of disfluent speech using phoneme-level modeling
Theodoros Kouzelis, Georgios Paraskevopoulos, Athanasios Katsamanis +1
The study of speech disorders can benefit greatly from time-aligned data. However, audio-text mismatches in disfluent speech cause rapid performance degradation for modern speech a…