Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures
Florian Rottach, William Rudman, Bastian Rieck +2
Studying how embeddings are organized in space not only enhances model interpretability but also uncovers factors that drive downstream task performance. In this paper, we present…
cs.LG2025
APP: Accelerated Path Patching with Task-Specific Pruning
Frauke Andersen, William Rudman, Ruochen Zhang +1
Circuit discovery is a key step in many mechanistic interpretability pipelines. Current methods, such as Path Patching, are computationally expensive and have limited in-depth circ…