3 papers
cs.SD2025
SeeingSounds: Learning Audio-to-Visual Alignment via Text
Simone Carnemolla, Matteo Pennisi, Chiara Russo +3
We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any…
cs.LG2024
Back to Supervision: Boosting Word Boundary Detection through Frame Classification
Simone Carnemolla, Salvatore Calcagno, Simone Palazzo +1
Speech segmentation at both word and phoneme levels is crucial for various speech processing tasks. It significantly aids in extracting meaningful units from an utterance, thus ena…
cs.CV2023
A baseline on continual learning methods for video action recognition
Giulia Castagnolo, Concetto Spampinato, Francesco Rundo +2
Continual learning has recently attracted attention from the research community, as it aims to solve long-standing limitations of classic supervisedly-trained models. However, most…