4 papers · 1 filter
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
Jiawei Zhang, Andrew Estornell, David D. Baek +2
Large Language Models (LLMs) exhibit strong but shallow alignment: they directly refuse harmful queries when a refusal is expected at the very start of an assistant turn, yet this…
Towards Understanding Distilled Reasoning Models: A Representational Approach
David D. Baek, Max Tegmark
In this paper, we investigate how model distillation impacts the development of reasoning features in large language models (LLMs). To explore this, we train a crosscoder on Qwen-s…
Harmonic Loss Trains Interpretable AI Models
David D. Baek, Ziming Liu, Riya Tyagi +1
In this paper, we introduce harmonic loss as an alternative supervisory signal for training neural networks and large language models (LLMs). Harmonic loss differs from standard cr…
Investigating Representation Universality: Case Study on Genealogical Representations
David D. Baek, Yuxiao Li, Max Tegmark
Motivated by interpretability and reliability, we investigate whether large language models (LLMs) deploy universal geometric structures to encode discrete, graph-structured knowle…