3 papers
cs.CL2026
MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines
Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang +5
This paper presents Murano, an open source framework for designing, running, and reproducing mechanistic interpretability studies of large language models, intended for researchers…
cs.CL2026
Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery
Alireza Bayat Makou, Jingcheng Niu, Subhabrata Dutta +1
Circuit discovery methods identify subgraphs that explain model behaviors, and structural differences between discovered circuits are commonly interpreted as evidence of distinct m…
cs.CL2024
LLM-based Rewriting of Inappropriate Argumentation using Reinforcement Learning from Machine Feedback
Timon Ziegenbein, Gabriella Skitalinskaya, Alireza Bayat Makou +1
Ensuring that online discussions are civil and productive is a major challenge for social media platforms. Such platforms usually rely both on users and on automated detection tool…