37 papers
LieSolver: PDE-Constrained Learning for IBVPs via Lie Symmetries
René P. Klausen, Ivan Timofeev, Jonas Naujoks +4
Initial-boundary value problems (IBVPs) provide the essential framework for modelling a wide range of phenomena in physics and engineering. We introduce a novel method for efficien…
The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network
Elias Sandmann, Sebastian Lapuschkin, Wojciech Samek
Recent mechanistic work has uncovered learned algorithms within neural networks, from modular arithmetic to search and planning in game-playing agents. But does algorithmic structu…
Predicting Future Behaviors in Reasoning Models Enables Better Steering
Evgenii Kortukov, Piotr Komorowski, Florian Klein +5
Deployed large reasoning models (LRMs) often behave unexpectedly. Test-time steering controls LRM outputs by intervening on their hidden representations, but it can degrade output…
Fast & Faithful Function Vectors
Minh An Pham, Anton Segeler, Thomas Wiegand +4
Function vectors (FVs) are task representations elicited during in-context learning that can be used to steer Large Language Models (LLMs). However, design choices in their formula…
PINNfluence: Interpreting PINNs through Influence Functions
Aleksander Krasowski, Jonas R. Naujoks, Moritz Weckbecker +5
Physics-informed neural networks (PINNs) have emerged as a powerful deep learning approach for solving partial differential equations (PDEs) in the physical sciences, yet their beh…
From Attribution to Action: A Human-Centered Application of Activation Steering
Tobias Labarta, Maximilian Dreyer, Katharina Weitz +2
Explainable AI (XAI) methods reveal which features influence model predictions, yet provide limited means for practitioners to act on these explanations. Activation steering of com…