3 citations · 3 across the 3 of their papers we have counts for
7 papers · 1 filter
Image-POSER: Reflective RL for Multi-Expert Image Generation and Editing
Hossein Mohebbi, Mohammed Abdulrahman, Yanting Miao +2
Recent advances in text-to-image generation have produced strong single-shot models, yet no individual system reliably executes the long, compositional prompts typical of creative…
FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models
Sankeerth Durvasula, Kavya Sreedhar, Zain Moustafa +6
Using diffusion transformers for media generation may require evaluating attention over extremely long sequences, with attention layers accounting for the majority of generation la…
Imagen 3
Imagen-Team-Google, :, Jason Baldridge +257
We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred…
Layered Diffusion Model for One-Shot High Resolution Text-to-Image Synthesis
Emaad Khwaja, Abdullah Rashwan, Ting Chen +3
We present a one-shot text-to-image diffusion model that can generate high-resolution images from natural language descriptions. Our model employs a layered U-Net architecture that…
Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning
Yanting Miao, William Loh, Suraj Kothawade +3
Text-to-image generative models have recently attracted considerable interest, enabling the synthesis of high-quality images from textual prompts. However, these models often lack…
Source-Free Domain Adaptation with Diffusion-Guided Source Data Generation
Shivang Chopra, Suraj Kothawade, Houda Aynaou +1
This paper introduces a novel approach to leverage the generalizability of Diffusion Models for Source-Free Domain Adaptation (DM-SFDA). Our proposed DMSFDA method involves fine-tu…