Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Image Generators are Generalist Vision Learners
Valentin Gabeur, Shangbang Long, Songyou Peng +22
Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language under…
cs.CV2024
Layered Diffusion Model for One-Shot High Resolution Text-to-Image Synthesis
Emaad Khwaja, Abdullah Rashwan, Ting Chen +3
We present a one-shot text-to-image diffusion model that can generate high-resolution images from natural language descriptions. Our model employs a layered U-Net architecture that…
cs.CV2024
Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
Cristina N. Vasconcelos, Abdullah Rashwan, Austin Waters +22
We address the long-standing problem of how to learn effective pixel-based image diffusion models at scale, introducing a remarkably simple greedy growing method for stable trainin…