1k citations · 1.2k across the 60 of their papers we have counts for
14 papers · 1 filter
Can Data Attribution Filter Out Subliminal Learning? Not Reliably
Moritz Weckbecker, Sweta Jena, Jonas Müller +5
Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data…
ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers
Jim Berend, Reduan Achtibat, Daniel Schäffer +4
Vision Transformers (ViTs) are central to most modern vision models, yet obtaining input attributions that are fine-grained, faithful, and stable remains challenging. Layer-wise Re…
Safety Signals to Verify NetOps Agents with Action-Level Granularity
Tobias Labarta, Frederik Pahde, Novak Boškov +6
Agentic Network Operations (NetOps) are an emerging paradigm promising to enable workload-aware, self-adjustable, and reliable autonomous networks. While agents have proven their v…
Concept-based explanation of gene expression prediction from H&E images
Amos Muench, Jonathan Thielmann, Reduan Achtibat +9
Recent advances in pathology foundation models have enabled accurate prediction of spatial transcriptomics (ST) from routine H&E images. However, existing explainability methods fo…
Predicting Future Behaviors in Reasoning Models Enables Better Steering
Evgenii Kortukov, Piotr Komorowski, Florian Klein +5
Deployed large reasoning models (LRMs) often behave unexpectedly. Test-time steering controls LRM outputs by intervening on their hidden representations, but it can degrade output…
Fast & Faithful Function Vectors
Minh An Pham, Anton Segeler, Thomas Wiegand +4
Function vectors (FVs) are task representations elicited during in-context learning that can be used to steer Large Language Models (LLMs). However, design choices in their formula…