6 papers
WaveMAE: Wavelet decomposition Masked Auto-Encoder for Remote Sensing
Vittorio Bernuzzi, Leonardo Rossi, Tomaso Fontanini +2
Self-supervised learning (SSL) has recently emerged as a key strategy for building foundation models in remote sensing, where the scarcity of annotated data limits the applicabilit…
SISMA: Semantic Face Image Synthesis with Mamba
Filippo Botti, Alex Ergasti, Tomaso Fontanini +3
Diffusion Models have become very popular for Semantic Image Synthesis (SIS) of human faces. Nevertheless, their training and inference is computationally expensive and their compu…
U-Shape Mamba: State Space Model for faster diffusion
Alex Ergasti, Filippo Botti, Tomaso Fontanini +3
Diffusion models have become the most popular approach for high-quality image generation, but their high computational cost still remains a significant challenge. To address this p…
FLAV: Rolling Flow matching for infinite Audio Video generation
Alex Ergasti, Giuseppe Gabriele Tarollo, Filippo Botti +4
Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to three critical requirements: quality of the generated samples, seamless multim…
Swin2-MoSE: A New Single Image Super-Resolution Model for Remote Sensing
Leonardo Rossi, Vittorio Bernuzzi, Tomaso Fontanini +2
Due to the limitations of current optical and sensor technologies and the high cost of updating them, the spectral and spatial resolution of satellites may not always meet desired…
MCGM: Mask Conditional Text-to-Image Generative Model
Rami Skaik, Leonardo Rossi, Tomaso Fontanini +1
Recent advancements in generative models have revolutionized the field of artificial intelligence, enabling the creation of highly-realistic and detailed images. In this study, we…