activity
20162026
most citedUnmasking Clever Hans Predictors and Assessing What Machines Really Learn

1k citations · 1.2k across the 60 of their papers we have counts for

collaborators
Showing 2026Show all

14 papers · 1 filter

cs.AI2026

Can Data Attribution Filter Out Subliminal Learning? Not Reliably

Moritz Weckbecker, Sweta Jena, Jonas Müller +5

Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data…

cs.CV2026

ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers

Jim Berend, Reduan Achtibat, Daniel Schäffer +4

Vision Transformers (ViTs) are central to most modern vision models, yet obtaining input attributions that are fine-grained, faithful, and stable remains challenging. Layer-wise Re…

cs.AI2026

Safety Signals to Verify NetOps Agents with Action-Level Granularity

Tobias Labarta, Frederik Pahde, Novak Boškov +6

Agentic Network Operations (NetOps) are an emerging paradigm promising to enable workload-aware, self-adjustable, and reliable autonomous networks. While agents have proven their v…

cs.CV2026

Concept-based explanation of gene expression prediction from H&E images

Amos Muench, Jonathan Thielmann, Reduan Achtibat +9

Recent advances in pathology foundation models have enabled accurate prediction of spatial transcriptomics (ST) from routine H&E images. However, existing explainability methods fo…

cs.LG2026

Predicting Future Behaviors in Reasoning Models Enables Better Steering

Evgenii Kortukov, Piotr Komorowski, Florian Klein +5

Deployed large reasoning models (LRMs) often behave unexpectedly. Test-time steering controls LRM outputs by intervening on their hidden representations, but it can degrade output…

cs.CL2026

Fast & Faithful Function Vectors

Minh An Pham, Anton Segeler, Thomas Wiegand +4

Function vectors (FVs) are task representations elicited during in-context learning that can be used to steer Large Language Models (LLMs). However, design choices in their formula…