13 papers
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
Nicolas Dufour, Lucas Degeorge, Arijit Ghosh +2
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator…
How far can we go with ImageNet for Text-to-Image generation?
L. Degeorge, A. Ghosh, N. Dufour +2
Recent text-to-image (T2I) generation models have achieved remarkable sucess by training on billion-scale datasets, following a `bigger is better' paradigm that prioritizes data qu…
Training-Free Synthetic Data Generation with Dual IP-Adapter Guidance
Luc Boudier, Loris Manganelli, Eleftherios Tsonis +2
Few-shot image classification remains challenging due to the limited availability of labeled examples. Recent approaches have explored generating synthetic training data using text…
DiO: Distilling Masked Diffusion Models into One-step Generator
Yuanzhi Zhu, Xi Wang, Stéphane Lathuilière +1
Masked Diffusion Models (MDMs) have emerged as a powerful generative modeling technique. Despite their remarkable results, they typically suffer from slow inference with several st…
Don't drop your samples! Coherence-aware training benefits Conditional diffusion
Nicolas Dufour, Victor Besnier, Vicky Kalogeiton +1
Conditional diffusion models are powerful generative models that can leverage various types of conditional information, such as class labels, segmentation masks, or text captions.…
Long Story Short: Story-level Video Understanding from 20K Short Films
Ridouane Ghermi, Xi Wang, Vicky Kalogeiton +1
Recent developments in vision-language models have significantly advanced video understanding. Existing datasets and tasks, however, have notable limitations. Most datasets are con…