12 citations · 12 across the 2 of their papers we have counts for
3 papers
cs.LG2025
Incoherence in Goal-Conditioned Autoregressive Models
Jacek Karwowski, Raymond Douglas
We investigate mathematically the notion of incoherence: a structural issue with reinforcement learning policies derived by naive goal-conditioning of autoregressive models. We foc…
cs.CY2025★ 12 cited
Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
Jan Kulveit, Raymond Douglas, Nora Ammann +3
This paper examines the systemic risks posed by incremental advancements in artificial intelligence, developing the concept of `gradual disempowerment', in contrast to the abrupt t…
cs.CL2024
Evaluating Language Model Character Traits
Francis Rhys Ward, Zejia Yang, Alex Jackson +7
Language models (LMs) can exhibit human-like behaviour, but it is unclear how to describe this behaviour without undue anthropomorphism. We formalise a behaviourist view of LM char…