2 papers
cs.AI2026
From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
Trilok Padhi, Ramneet Kaur, Krishiv Agarwal +9
Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive environments. Despite their growing capabi…
cs.CR2026
Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
Krishiv Agarwal, Ramneet Kaur, Colin Samplawski +6
Effective safety auditing of large language models (LLMs) demands tools that go beyond black-box probing and systematically uncover vulnerabilities rooted in model internals. We pr…