activity
20242026
most citedMechanistic understanding and validation of large AI models with SemanticLens

2 citations · 2 across the 3 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG2025

LieSolver: PDE-Constrained Learning for IBVPs via Lie Symmetries

René P. Klausen, Ivan Timofeev, Jonas Naujoks +4

Initial-boundary value problems (IBVPs) provide the essential framework for modelling a wide range of phenomena in physics and engineering. We introduce a novel method for efficien…

cs.LG2025

Attribution-Guided Decoding

Piotr Komorowski, Elena Golimblevskaia, Reduan Achtibat +3

The capacity of Large Language Models (LLMs) to follow complex instructions and generate factually accurate text is critical for their real-world application. However, standard dec…

cs.LG2025

Leveraging Influence Functions for Resampling Data in Physics-Informed Neural Networks

Jonas R. Naujoks, Aleksander Krasowski, Moritz Weckbecker +5

Physics-informed neural networks (PINNs) offer a powerful approach to solving partial differential equations (PDEs), which are ubiquitous in the quantitative sciences. Applied to b…

cs.LG2025

Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs

Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Reduan Achtibat +5

Large Language Models (LLMs) are widely deployed in real-world applications, yet their internal mechanisms remain difficult to interpret and control, limiting our ability to diagno…

cs.LG2025

From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance

Maximilian Dreyer, Lorenz Hufe, Jim Berend +3

Transformer-based CLIP models are widely used for text-image probing and feature extraction, making it relevant to understand the internal mechanisms behind their predictions. Whil…

cs.LG2025

FADE: Why Bad Descriptions Happen to Good Features

Bruno Puri, Aakriti Jain, Elena Golimblevskaia +4

Recent advances in mechanistic interpretability have highlighted the potential of automating interpretability pipelines in analyzing the latent representations within LLMs. While t…