53 citations · 54 across the 2 of their papers we have counts for
2 papers
cs.CL2023★ 53 cited
The Capacity for Moral Self-Correction in Large Language Models
Deep Ganguli, Amanda Askell, Nicholas Schiefer +46
We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmf…
hep-ph2022★ 1 cited
Dark matter scattering in astrophysical media: collective effects
William DeRocco, Marios Galanis, Robert Lasenby
It is well-known that stars have the potential to be excellent dark matter detectors. Infalling dark matter that scatters within stars could lead to a range of observational signat…