2 papers
cs.AI2026
Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
Marco Valentino, Geonhee Kim, Dhairya Dalal +2
Large language models (LLMs) exhibit reasoning biases, often conflating content plausibility with formal logical validity. This can lead to wrong inferences in critical domains, wh…
cs.CL2025
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
Geonhee Kim, Marco Valentino, André Freitas
Recent studies on reasoning in language models (LMs) have sparked a debate on whether they can learn systematic inferential principles or merely exploit superficial patterns in the…