170 citations · 523 across the 13 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023★ 31 cited
Towards Evaluating AI Systems for Moral Status Using Self-Reports
Ethan Perez, Robert Long
As AI systems become more advanced and widely deployed, there will likely be increasing debate over whether AI systems could have conscious experiences, desires, or other states of…
cs.LG2023★ 26 cited
Studying Large Language Model Generalization with Influence Functions
Roger Grosse, Juhan Bae, Cem Anil +14
When trying to gain better visibility into a machine learning model in order to understand and mitigate the associated risks, a potentially valuable source of evidence is: which tr…