Showing 2026 · cs.CLShow all
2 papers · 2 filters
cs.CL2026
Task-Specific Knowledge Distillation via Intermediate Probes
Ryan Brown, Chris Russell
Knowledge distillation from large language models (LLMs) assumes that the teacher's output distribution is a high-quality training signal. On reasoning tasks, this assumption is fr…
cs.CL2026
LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
William Lugoloobi, Thomas Foster, William Bankes +1
Running LLMs with extended reasoning on every problem is expensive, but determining which inputs actually require additional compute remains challenging. We investigate whether the…