3 papers
cs.LG2026
The Double Life of Code World Models: Provably Unmasking Malicious Behavior Through Execution Traces
Subramanyam Sahoo
Large language models (LLMs) increasingly generate code with minimal human oversight, raising critical concerns about backdoor injection and malicious behavior. We present Cross-Tr…
cs.CV2025
The Deepfake Detective: Interpreting Neural Forensics Through Sparse Features and Manifolds
Subramanyam Sahoo, Jared Junkin
Deepfake detection models have achieved high accuracy in identifying synthetic media, but their decision processes remain largely opaque. In this paper we present a mechanistic int…
cs.LG2025
The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems
Subramanyam Sahoo, Jared Junkin
Embodied AI agents exploit reward signal flaws through reward hacking, achieving high proxy scores while failing true objectives. We introduce Mechanistically Interpretable Task De…