4 citations · 7 across the 7 of their papers we have counts for
6 papers · 1 filter
ALOHa: A New Measure for Hallucination in Captioning Models
Suzanne Petryk, David M. Chan, Anish Kachinthaya +4
Despite recent advances in multimodal pre-training for visual description, state-of-the-art models still produce captions containing errors, such as hallucinating objects not prese…
CLAIR: Evaluating Image Captions with Large Language Models
David Chan, Suzanne Petryk, Joseph E. Gonzalez +2
The evaluation of machine-generated image captions poses an interesting yet persistent challenge. Effective evaluation measures must consider numerous dimensions of similarity, inc…
Simple Token-Level Confidence Improves Caption Correctness
Suzanne Petryk, Spencer Whitehead, Joseph E. Gonzalez +3
The ability to judge whether a caption correctly describes an image is a critical part of vision-language understanding. However, state-of-the-art models often misinterpret the cor…
Prior Knowledge-Guided Attention in Self-Supervised Vision Transformers
Kevin Miao, Akash Gokul, Raghav Singh +5
Recent trends in self-supervised representation learning have focused on removing inductive biases from training pipelines. However, inductive biases can be useful in settings when…
On Guiding Visual Attention with Language Specification
Suzanne Petryk, Lisa Dunlap, Keyan Nasseri +3
While real world challenges typically define visual categories with language words or phrases, most visual classification methods define categories with numerical indices. However,…
NBDT: Neural-Backed Decision Trees
Alvin Wan, Lisa Dunlap, Daniel Ho +6
Machine learning applications such as finance and medicine demand accurate and justifiable predictions, barring most deep learning methods from use. In response, previous work comb…