2 papers
cs.CL2026
Perfect Detection, Failed Control: The Geometry of Knowing vs. Steering in Language Models
Cosimo Galeone, Anna Ettorre, Minsu Park +2
A central aspiration of mechanistic interpretability is controllability: if we know where a behavior is represented in a model's activations, we should be able to modify it. This r…
cs.CL2026
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
Cosimo Galeone, Minsu Park, Giuseppe Ettorre +1
Deployed language models must produce outputs that are both correct and format-compliant. We study this structured-output reliability gap using two mathematical benchmarks -- GSM8K…