most citedDavidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

8 citations · 13 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG2024★ 1 cited

Evaluating Numerical Reasoning in Text-to-Image Models

Ivana Kajić, Olivia Wiles, Isabela Albuquerque +4

Text-to-image generative models are capable of producing high-quality images that often faithfully depict concepts described using natural language. In this work, we comprehensivel…

cs.CV2024

Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models

Cristina N. Vasconcelos, Abdullah Rashwan, Austin Waters +22

We address the long-standing problem of how to learn effective pixel-based image diffusion models at scale, introducing a remarkably simple greedy growing method for stable trainin…

cs.CV2024★ 1 cited

DOCCI: Descriptions of Connected and Contrasting Images

Yasumasa Onoe, Sunayana Rane, Zachary Berger +9

Vision-language datasets are vital for both text-to-image (T2I) and image-to-text (I2T) research. However, current datasets lack descriptions with fine-grained detail that would al…

cs.CV2024

Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Olivia Wiles, Chuhan Zhang, Isabela Albuquerque +11

While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While previous work has evaluated T2I al…

cs.CV2023★ 3 cited

Rich Human Feedback for Text-to-Image Generation

Youwei Liang, Junfeng He, Gang Li +15

Recent Text-to-Image (T2I) generation models such as Stable Diffusion and Imagen have made significant progress in generating high-resolution images based on text descriptions. How…

cs.CV2023★ 8 cited

Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

Jaemin Cho, Yushi Hu, Roopal Garg +6

Evaluating text-to-image models is notoriously difficult. A strong recent approach for assessing text-image faithfulness is based on QG/A (question generation and answering), which…