124 citations · 125 across the 2 of their papers we have counts for
2 papers
cs.CY2023★ 124 cited
Harms from Increasingly Agentic Algorithmic Systems
Alan Chan, Rebecca Salganik, Alva Markelius +19
Research in Fairness, Accountability, Transparency, and Ethics (FATE) has established many sources and forms of algorithmic harm, in domains as diverse as health care, finance, pol…
cs.LG2023★ 1 cited
On The Fragility of Learned Reward Functions
Lev McKinney, Yawen Duan, David Krueger +1
Reward functions are notoriously difficult to specify, especially for tasks with complex goals. Reward learning approaches attempt to infer reward functions from human feedback and…