4 papers · 1 filter
AsyncPatch Diffusion: spatially-flexible image generation
Samuele Papa, Valentin De Bortoli, Guillaume Couairon +3
Standard diffusion models corrupt an entire sample with a single shared noise level, forcing all spatial regions to follow the same denoising trajectory. We introduce AsyncPatch Di…
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
Amir Mohammad Karimi Mamaghan, Samuele Papa, Karl Henrik Johansson +2
Object-centric (OC) representations, which model visual scenes as compositions of discrete objects, have the potential to be used in various downstream tasks to achieve systematic…
NARAIM: Native Aspect Ratio Autoregressive Image Models
Daniel Gallo Fernández, Robert van der Klis, RÄzvan-Andrei MatiÅan +4
While vision transformers are able to solve a wide variety of computer vision tasks, no pre-training method has yet demonstrated the same scaling laws as observed in language model…
How to Train Neural Field Representations: A Comprehensive Study and Benchmark
Samuele Papa, Riccardo Valperga, David Knigge +4
Neural fields (NeFs) have recently emerged as a versatile method for modeling signals of various modalities, including images, shapes, and scenes. Subsequently, a number of works h…