22 papers · 1 filter
LieSolver: PDE-Constrained Learning for IBVPs via Lie Symmetries
René P. Klausen, Ivan Timofeev, Jonas Naujoks +4
Initial-boundary value problems (IBVPs) provide the essential framework for modelling a wide range of phenomena in physics and engineering. We introduce a novel method for efficien…
The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network
Elias Sandmann, Sebastian Lapuschkin, Wojciech Samek
Recent mechanistic work has uncovered learned algorithms within neural networks, from modular arithmetic to search and planning in game-playing agents. But does algorithmic structu…
Predicting Future Behaviors in Reasoning Models Enables Better Steering
Evgenii Kortukov, Piotr Komorowski, Florian Klein +5
Deployed large reasoning models (LRMs) often behave unexpectedly. Test-time steering controls LRM outputs by intervening on their hidden representations, but it can degrade output…
PINNfluence: Interpreting PINNs through Influence Functions
Aleksander Krasowski, Jonas R. Naujoks, Moritz Weckbecker +5
Physics-informed neural networks (PINNs) have emerged as a powerful deep learning approach for solving partial differential equations (PDEs) in the physical sciences, yet their beh…
Playing the network backward: A Game Theoretic Attribution Framework
Jakob Paul Zimmermann, Jim Berend, Georg Loho +2
Attribution methods explain which input features drive a model's prediction, making them central to model debugging and mechanistic interpretability. Yet backward attribution metho…
Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Reduan Achtibat +5
Large Language Models (LLMs) are widely deployed in real-world applications, yet their internal mechanisms remain difficult to interpret and control, limiting our ability to diagno…