3 papers
cs.CV2025
SISMA: Semantic Face Image Synthesis with Mamba
Filippo Botti, Alex Ergasti, Tomaso Fontanini +3
Diffusion Models have become very popular for Semantic Image Synthesis (SIS) of human faces. Nevertheless, their training and inference is computationally expensive and their compu…
cs.CV2025
U-Shape Mamba: State Space Model for faster diffusion
Alex Ergasti, Filippo Botti, Tomaso Fontanini +3
Diffusion models have become the most popular approach for high-quality image generation, but their high computational cost still remains a significant challenge. To address this p…
cs.CV2025
FLAV: Rolling Flow matching for infinite Audio Video generation
Alex Ergasti, Giuseppe Gabriele Tarollo, Filippo Botti +4
Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to three critical requirements: quality of the generated samples, seamless multim…