4 papers
Incoherence in Goal-Conditioned Autoregressive Models
Jacek Karwowski, Raymond Douglas
We investigate mathematically the notion of incoherence: a structural issue with reinforcement learning policies derived by naive goal-conditioning of autoregressive models. We foc…
Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
Jan Kulveit, Raymond Douglas, Nora Ammann +3
This paper examines the systemic risks posed by incremental advancements in artificial intelligence, developing the concept of `gradual disempowerment', in contrast to the abrupt t…
Mitigating the Influence of Distractor Tasks in LMs with Prior-Aware Decoding
Raymond Douglas, Andis Draguns, Tomáš GavenÄiak
The broad capabilities of Language Models (LMs) can be limited by their sensitivity to distractor tasks: LMs can infer secondary tasks from the prompt in addition to the intended o…
Evaluating Language Model Character Traits
Francis Rhys Ward, Zejia Yang, Alex Jackson +7
Language models (LMs) can exhibit human-like behaviour, but it is unclear how to describe this behaviour without undue anthropomorphism. We formalise a behaviourist view of LM char…