3 papers
cs.LG2026
Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit
Sai Adith Senthil Kumar
One approach to mechanistic interpretability explains behavior through circuits: the components and connections that carry it. Frozen discovery often returns hundreds of edges, mak…
cs.CL2026
When Built-in Thinking Helps and Hurts: Constraint-Level Error Shifts in Instruction Following
Sai Adith Senthil Kumar
Large reasoning models (LRMs) often improve math and coding performance, but their effect on instruction following is unclear. We study IFEval with Qwen3 models (1.7B-32B), using s…
cs.CL2025
Can LLMs Simulate Personas with Reversed Performance? A Systematic Investigation for Counterfactual Instruction Following in Math Reasoning Context
Sai Adith Senthil Kumar, Hao Yan, Saipavan Perepa +2
Large Language Models (LLMs) are now increasingly widely used to simulate personas in virtual environments, leveraging their instruction-following capability. However, we discovere…