827 citations · 1.6k across the 9 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2022★ 133 cited
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz +31
We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines. We use reinforcement lear…
cs.LG2021★ 1 cited
Statistical discrimination in learning agents
Edgar A. Duéñez-Guzmán, Kevin R. McKee, Yiran Mao +9
Undesired bias afflicts both human and algorithmic decision making, and may be especially prevalent when information processing trade-offs incentivize the use of heuristics. One pr…