From the 1 of 9 linked papers with an AI index.
9 papers
Constitutional Midtraining: Content Presence Drives Alignment Gains
Desiree Cho, Cameron Tice, Bernie Hogan +4
The paper investigates inserting constitutionally‑derived content during midtraining of large language models to improve the durability of alignment, showing reduced blackmail tend…
Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks
Nathaniel Mitrani Hadida, Sassan Bhanji, Cameron Tice +1
Chain-of-thought (CoT) reasoning provides a significant performance uplift to LLMs by enabling planning, exploration, and deliberation of their actions. CoT is also a powerful tool…
A transformer architecture alteration to incentivise externalised reasoning
Elizabeth Pavlova, Mariia Koroliuk, Karthik Viswanathan +3
We propose a new architectural change, and post-training pipeline, for making LLMs more verbose reasoners by teaching a model to truncate forward passes early. We augment an existi…
Setting up for failure: automatic discovery of the neural mechanisms of cognitive errors
Puria Radmard, Paul M. Bays, Máté Lengyel
Discovering the neural mechanisms underpinning cognition is one of the grand challenges of neuroscience. However, previous approaches for building models of RNN dynamics that expla…
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
Cameron Tice, Puria Radmard, Samuel Ratnam +3
Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descri…
Diagnosing Pathological Chain-of-Thought in Reasoning Models
Manqing Liu, David Williams-King, Ida Caspary +5
Chain-of-thought (CoT) reasoning is fundamental to modern LLM architectures and represents a critical intervention point for AI safety. However, CoT reasoning may exhibit failure m…