440 citations · 1.1k across the 16 of their papers we have counts for
4 papers · 1 filter
Self-correcting LLM-controlled Diffusion Models
Tsung-Han Wu, Long Lian, Joseph E. Gonzalez +2
Text-to-image generation has witnessed significant progress with the advent of diffusion models. Despite the ability to generate photorealistic images, current text-to-image diffus…
CLAIR: Evaluating Image Captions with Large Language Models
David Chan, Suzanne Petryk, Joseph E. Gonzalez +2
The evaluation of machine-generated image captions poses an interesting yet persistent challenge. Effective evaluation measures must consider numerous dimensions of similarity, inc…
Simple Token-Level Confidence Improves Caption Correctness
Suzanne Petryk, Spencer Whitehead, Joseph E. Gonzalez +3
The ability to judge whether a caption correctly describes an image is a critical part of vision-language understanding. However, state-of-the-art models often misinterpret the cor…
Context-Aware Streaming Perception in Dynamic Environments
Gur-Eyal Sela, Ionel Gog, Justin Wong +9
Efficient vision works maximize accuracy under a latency budget. These works evaluate accuracy offline, one image at a time. However, real-time vision applications like autonomous…