3 papers
cs.LG2026
Training on Documents About Monitoring Leads to CoT Obfuscation
Reilly Haskins, Bilal Chughtai, Joshua Engels
Chain-of-thought (CoT) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models faithfully externalizing their…
cs.LG2026
Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
Reilly Haskins, Benjamin Adams
Knowledge distillation compresses a larger neural model (teacher) into smaller, faster student models by training them to match teacher outputs. However, the internal computational…
cs.LG2025
KEA Explain: Explanations of Hallucinations using Graph Kernel Analysis
Reilly Haskins, Benjamin Adams
Large Language Models (LLMs) frequently generate hallucinations: statements that are syntactically plausible but lack factual grounding. This research presents KEA (Kernel-Enriched…