1 citations · 1 across the 2 of their papers we have counts for
3 papers · 1 filter
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
Benjamin Feuer, Micah Goldblum, Teresa Datta +5
The release of ChatGPT in November 2022 sparked an explosion of interest in post-training and an avalanche of new preference optimization (PO) methods. These methods claim superior…
Reckoning with the Disagreement Problem: Explanation Consensus as a Training Objective
Avi Schwarzschild, Max Cembalest, Karthik Rao +2
As neural networks increasingly make critical decisions in high-stakes settings, monitoring and explaining their behavior in an understandable and trustworthy manner is a necessity…
Tensions Between the Proxies of Human Values in AI
Teresa Datta, Daniel Nissani, Max Cembalest +3
Motivated by mitigating potentially harmful impacts of technologies, the AI community has formulated and accepted mathematical definitions for certain pillars of accountability: e.…