2 papers
cs.LG2026
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
Jinghan Jia, Joe Benton, Eric Easley
Chain-of-thought (CoT) reasoning is useful for monitoring language models only when the reasoning trace faithfully reflects the computation that produces the final answer. However,…
cs.LG2026
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
Eric Easley, Sebastian Farquhar
We address jailbreaks, backdoors, and unlearning for large language models (LLMs). Unlike prior work, which trains LLMs based on their actions when given malign instructions, our m…