collaborators

6 papers

cs.CV2025

WaveMAE: Wavelet decomposition Masked Auto-Encoder for Remote Sensing

Vittorio Bernuzzi, Leonardo Rossi, Tomaso Fontanini +2

Self-supervised learning (SSL) has recently emerged as a key strategy for building foundation models in remote sensing, where the scarcity of annotated data limits the applicabilit…

cs.CV2025

SISMA: Semantic Face Image Synthesis with Mamba

Filippo Botti, Alex Ergasti, Tomaso Fontanini +3

Diffusion Models have become very popular for Semantic Image Synthesis (SIS) of human faces. Nevertheless, their training and inference is computationally expensive and their compu…

cs.CV2025

U-Shape Mamba: State Space Model for faster diffusion

Alex Ergasti, Filippo Botti, Tomaso Fontanini +3

Diffusion models have become the most popular approach for high-quality image generation, but their high computational cost still remains a significant challenge. To address this p…

cs.CV2025

FLAV: Rolling Flow matching for infinite Audio Video generation

Alex Ergasti, Giuseppe Gabriele Tarollo, Filippo Botti +4

Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to three critical requirements: quality of the generated samples, seamless multim…

cs.CV2024

Swin2-MoSE: A New Single Image Super-Resolution Model for Remote Sensing

Leonardo Rossi, Vittorio Bernuzzi, Tomaso Fontanini +2

Due to the limitations of current optical and sensor technologies and the high cost of updating them, the spectral and spatial resolution of satellites may not always meet desired…

cs.CV2024

MCGM: Mask Conditional Text-to-Image Generative Model

Rami Skaik, Leonardo Rossi, Tomaso Fontanini +1

Recent advancements in generative models have revolutionized the field of artificial intelligence, enabling the creation of highly-realistic and detailed images. In this study, we…