3 papers
cs.LG2026
Inverted Detection and Control in Steering Vectors
Max Torop, Aria Masoomi, Jennifer Dy
Steering vectors (SVs) are widely used to influence the expression of concepts (e.g., truthfulness) in large language model outputs. A key assumption underpinning SVs is that they…
cs.LG2025
OrdShap: Feature Position Importance for Sequential Black-Box Models
Davin Hill, Brian L. Hill, Aria Masoomi +3
Sequential deep learning models excel in domains with temporal or sequential dependencies, but their complexity necessitates post-hoc feature attribution methods for understanding…
cs.LG2025
Axiomatic Explainer Globalness via Optimal Transport
Davin Hill, Josh Bone, Aria Masoomi +2
Explainability methods are often challenging to evaluate and compare. With a multitude of explainers available, practitioners must often compare and select explainers based on quan…