1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 1 cited
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
Benjamin Feuer, Micah Goldblum, Teresa Datta +5
The release of ChatGPT in November 2022 sparked an explosion of interest in post-training and an avalanche of new preference optimization (PO) methods. These methods claim superior…
cs.CY2024
Strategies for Increasing Corporate Responsible AI Prioritization
Angelina Wang, Teresa Datta, John P. Dickerson
Responsible artificial intelligence (RAI) is increasingly recognized as a critical concern. However, the level of corporate RAI prioritization has not kept pace. In this work, we c…