activity
20242026
most citedGemma 4 Technical Report

1 citations · 1 across the 2 of their papers we have counts for

collaborators

6 papers

cs.CL20261 cited

Gemma 4 Technical Report

Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…

cs.CV2026

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models

Sankeerth Durvasula, Kavya Sreedhar, Zain Moustafa +6

Using diffusion transformers for media generation may require evaluating attention over extremely long sequences, with attention layers accounting for the majority of generation la…

cs.CV2025

Image-POSER: Reflective RL for Multi-Expert Image Generation and Editing

Hossein Mohebbi, Mohammed Abdulrahman, Yanting Miao +2

Recent advances in text-to-image generation have produced strong single-shot models, yet no individual system reliably executes the long, compositional prompts typical of creative…

cs.LG2025

A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models

Yanting Miao, William Loh, Pacal Poupart +1

Recent work uses reinforcement learning (RL) to fine-tune text-to-image diffusion models, improving text-image alignment and sample quality. However, existing approaches introduce…

cs.CV2024

Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning

Yanting Miao, William Loh, Suraj Kothawade +3

Text-to-image generative models have recently attracted considerable interest, enabling the synthesis of high-quality images from textual prompts. However, these models often lack…

cs.CV2024

Imagen 3

Imagen-Team-Google, :, Jason Baldridge +257

We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred…