2 papers
cs.LG2026
Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior
Christopher J. Anders, Henrique Da Silva Gameiro, Nico Daheim +1
One way to understand LLM behavior is to trace its output back to the training data. Two types of measures are commonly used for output tracing: data-similarity and data-influence.…
cs.CV2025
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
Frederik Pahde, Maximilian Dreyer, Leander Weber +5
With a growing interest in understanding neural network prediction strategies, Concept Activation Vectors (CAVs) have emerged as a popular tool for modeling human-understandable co…