Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes
Dylan Jayabahu
A truth probe fitted where truthful reporting and a task's prescribed action coincide cannot distinguish those targets from its fitting labels alone. We call this failure of semant…
cs.LG2026
The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning
Dylan Jayabahu, Tinuade Adeleke
Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to…