From the 1 of 15 linked papers with an AI index.
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
Florian Eichin, Yupei Du, Philipp Mondorf +3
Post-hoc interpretability methods typically attribute a model's behavior to its components, data, or training trajectory in isolation, and are often tied to a particular level of g…
cs.LG2026
Tracing Uncertainty in Language Model "Reasoning"
Nils Grünefeld, Bertram Højer, Philipp Mondorf +5
Language model (LM) "reasoning", commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics underlying this process remain…
cs.LG2025
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
Philipp Mondorf, Sondre Wold, Barbara Plank
A fundamental question in interpretability research is to what extent neural networks, particularly language models, implement reusable functions through subnetworks that can be co…