collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

WaveMAE: Wavelet decomposition Masked Auto-Encoder for Remote Sensing

Vittorio Bernuzzi, Leonardo Rossi, Tomaso Fontanini +2

Self-supervised learning (SSL) has recently emerged as a key strategy for building foundation models in remote sensing, where the scarcity of annotated data limits the applicabilit…

cs.CV2025

SISMA: Semantic Face Image Synthesis with Mamba

Filippo Botti, Alex Ergasti, Tomaso Fontanini +3

Diffusion Models have become very popular for Semantic Image Synthesis (SIS) of human faces. Nevertheless, their training and inference is computationally expensive and their compu…

cs.CV2025

U-Shape Mamba: State Space Model for faster diffusion

Alex Ergasti, Filippo Botti, Tomaso Fontanini +3

Diffusion models have become the most popular approach for high-quality image generation, but their high computational cost still remains a significant challenge. To address this p…

cs.CV2025

FLAV: Rolling Flow matching for infinite Audio Video generation

Alex Ergasti, Giuseppe Gabriele Tarollo, Filippo Botti +4

Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to three critical requirements: quality of the generated samples, seamless multim…

cs.CV2024

Mamba-ST: State Space Model for Efficient Style Transfer

Filippo Botti, Alex Ergasti, Leonardo Rossi +4

The goal of style transfer is, given a content image and a style source, generating a new image preserving the content but with the artistic representation of the style source. Mos…