1 citations · 1 across the 2 of their papers we have counts for
6 papers
Gemma 4 Technical Report
Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…
FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models
Sankeerth Durvasula, Kavya Sreedhar, Zain Moustafa +6
Using diffusion transformers for media generation may require evaluating attention over extremely long sequences, with attention layers accounting for the majority of generation la…
Image-POSER: Reflective RL for Multi-Expert Image Generation and Editing
Hossein Mohebbi, Mohammed Abdulrahman, Yanting Miao +2
Recent advances in text-to-image generation have produced strong single-shot models, yet no individual system reliably executes the long, compositional prompts typical of creative…
A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models
Yanting Miao, William Loh, Pacal Poupart +1
Recent work uses reinforcement learning (RL) to fine-tune text-to-image diffusion models, improving text-image alignment and sample quality. However, existing approaches introduce…
Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning
Yanting Miao, William Loh, Suraj Kothawade +3
Text-to-image generative models have recently attracted considerable interest, enabling the synthesis of high-quality images from textual prompts. However, these models often lack…
Imagen 3
Imagen-Team-Google, :, Jason Baldridge +257
We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred…