26 citations · 60 across the 14 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
Benjamin Feuer, Micah Goldblum, Teresa Datta +5
The release of ChatGPT in November 2022 sparked an explosion of interest in post-training and an avalanche of new preference optimization (PO) methods. These methods claim superior…
cs.LG2020
Fairness Through Robustness: Investigating Robustness Disparity in Deep Learning
Vedant Nanda, Samuel Dooley, Sahil Singla +2
Deep neural networks (DNNs) are increasingly used in real-world applications (e.g. facial recognition). This has resulted in concerns about the fairness of decisions made by these…