Showing cs.AIShow all
2 papers · 1 filter
cs.AI2023
Discovering Variable Binding Circuitry with Desiderata
Xander Davies, Max Nadeau, Nikhil Prakash +2
Recent work has shown that computation in language models may be human-understandable, with successful efforts to localize and intervene on both single-unit features and input-outp…
cs.AI2023
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Stephen Casper, Xander Davies, Claudia Shi +29
Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of…