activity
20192026
most citedExplanations of Black-Box Models based on Directional Feature Interactions

9 citations · 19 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

11 papers · 1 filter

cs.LG2026

Inverted Detection and Control in Steering Vectors

Max Torop, Aria Masoomi, Jennifer Dy

Steering vectors (SVs) are widely used to influence the expression of concepts (e.g., truthfulness) in large language model outputs. A key assumption underpinning SVs is that they…

cs.LG2026

Geometric Signatures of Reasoning: A Spectral Perspective on Task Hardness

Aria Masoomi, Mahsa Bazzaz, Adel Javanmard +1

Chain-of-thought (CoT) reasoning enables large language models (LLMs) to solve complex problems by generating intermediate reasoning steps. While much attention has been paid to th…

cs.LG2025

DISCO: Disentangled Communication Steering for Large Language Models

Max Torop, Aria Masoomi, Masih Eskandar +1

A variety of recent methods guide large language model outputs via the inference-time addition of steering vectors to residual-stream or attention-head representations. In contrast…

cs.LG2025

OrdShap: Feature Position Importance for Sequential Black-Box Models

Davin Hill, Brian L. Hill, Aria Masoomi +3

Sequential deep learning models excel in domains with temporal or sequential dependencies, but their complexity necessitates post-hoc feature attribution methods for understanding…

cs.LG2024

Axiomatic Explainer Globalness via Optimal Transport

Davin Hill, Josh Bone, Aria Masoomi +2

Explainability methods are often challenging to evaluate and compare. With a multitude of explainers available, practitioners must often compare and select explainers based on quan…

cs.LG2023

SmoothHess: ReLU Network Feature Interactions via Stein's Lemma

Max Torop, Aria Masoomi, Davin Hill +3

Several recent methods for interpretability model feature interactions by looking at the Hessian of a neural network. This poses a challenge for ReLU networks, which are piecewise-…