9 citations · 19 across the 10 of their papers we have counts for
11 papers · 1 filter
Inverted Detection and Control in Steering Vectors
Max Torop, Aria Masoomi, Jennifer Dy
Steering vectors (SVs) are widely used to influence the expression of concepts (e.g., truthfulness) in large language model outputs. A key assumption underpinning SVs is that they…
Geometric Signatures of Reasoning: A Spectral Perspective on Task Hardness
Aria Masoomi, Mahsa Bazzaz, Adel Javanmard +1
Chain-of-thought (CoT) reasoning enables large language models (LLMs) to solve complex problems by generating intermediate reasoning steps. While much attention has been paid to th…
DISCO: Disentangled Communication Steering for Large Language Models
Max Torop, Aria Masoomi, Masih Eskandar +1
A variety of recent methods guide large language model outputs via the inference-time addition of steering vectors to residual-stream or attention-head representations. In contrast…
OrdShap: Feature Position Importance for Sequential Black-Box Models
Davin Hill, Brian L. Hill, Aria Masoomi +3
Sequential deep learning models excel in domains with temporal or sequential dependencies, but their complexity necessitates post-hoc feature attribution methods for understanding…
Axiomatic Explainer Globalness via Optimal Transport
Davin Hill, Josh Bone, Aria Masoomi +2
Explainability methods are often challenging to evaluate and compare. With a multitude of explainers available, practitioners must often compare and select explainers based on quan…
SmoothHess: ReLU Network Feature Interactions via Stein's Lemma
Max Torop, Aria Masoomi, Davin Hill +3
Several recent methods for interpretability model feature interactions by looking at the Hessian of a neural network. This poses a challenge for ReLU networks, which are piecewise-…