7 papers
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes
Alec Harris, Kasey Corra, Archie Chaudhury +1
Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. H…
ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modeling
Archie Chaudhury
Contemporary language models are dominated by the transformer architecture, which leverages self-attention mechanisms to enable more efficient, parallelized training across a wide…
Forgetting is Not Erasure: Recovering Latent Knowledge via Transport Keys
Archie Chaudhury
Catastrophic forgetting is often framed as a representational problem: after sequential training, a model appears to lose the features that supported performance on earlier tasks.…
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
Leon Eshuijs, Archie Chaudhury, Alan McBeth +1
LLM-as-a-judge is widely used as a scalable substitute for human evaluation, yet current approaches rely on black-box access and struggle to detect subtle dishonesty, such as sycop…
The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders
Shikhar Shiromani, Archie Chaudhury, Sri Pranav Kunda
Large Language Models (LLMs) frequently exhibit unfaithful behavior, producing a final answer that differs significantly from their internal chain of thought (CoT) reasoning in ord…
Alignment is Localized: A Causal Probe into Preference Layers
Archie Chaudhury
Reinforcement Learning frameworks, particularly those utilizing human annotations, have become an increasingly popular method for preference fine-tuning, where the outputs of a lan…