175 citations · 372 across the 14 of their papers we have counts for
13 papers · 1 filter
Bias at the End of the Score
Salma Abdel Magid, Grace Guo, Esin Tureci +4
Reward models (RMs) are inherently non-neutral value functions designed and trained to encode specific objectives, such as human preferences or text-image alignment. RMs have becom…
Attention IoU: Examining Biases in CelebA using Attention Maps
Aaron Serianni, Tyler Zhu, Olga Russakovsky +1
Computer vision models have been shown to exhibit and amplify biases across a wide array of datasets and tasks. Existing methods for quantifying bias in classification models prima…
UFO: A unified method for controlling Understandability and Faithfulness Objectives in concept-based explanations for CNNs
Vikram V. Ramaswamy, Sunnie S. Y. Kim, Ruth Fong +1
Concept-based explanations for convolutional neural networks (CNNs) aim to explain model behavior and outputs using a pre-defined set of semantic concepts (e.g., the model recogniz…
Overwriting Pretrained Bias with Finetuning Data
Angelina Wang, Olga Russakovsky
Transfer learning is beneficial by allowing the expressive features of models pretrained on large-scale datasets to be finetuned for the target task of smaller, more domain-specifi…
GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition
Vikram V. Ramaswamy, Sing Yu Lin, Dora Zhao +4
Current dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce st…
Overlooked factors in concept-based explanations: Dataset choice, concept learnability, and human capability
Vikram V. Ramaswamy, Sunnie S. Y. Kim, Ruth Fong +1
Concept-based interpretability methods aim to explain deep neural network model predictions using a predefined set of semantic concepts. These methods evaluate a trained model on a…