4 papers
Inverted Detection and Control in Steering Vectors
Max Torop, Aria Masoomi, Jennifer Dy
Steering vectors (SVs) are widely used to influence the expression of concepts (e.g., truthfulness) in large language model outputs. A key assumption underpinning SVs is that they…
DISCO: Disentangled Communication Steering for Large Language Models
Max Torop, Aria Masoomi, Masih Eskandar +1
A variety of recent methods guide large language model outputs via the inference-time addition of steering vectors to residual-stream or attention-head representations. In contrast…
Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study
Max Torop, Masih Eskandar, Nicholas Kurtansky +6
Artificial Intelligence models have demonstrated significant success in diagnosing skin diseases, including cancer, showing the potential to assist clinicians in their analysis. Ho…
Axiomatic Explainer Globalness via Optimal Transport
Davin Hill, Josh Bone, Aria Masoomi +2
Explainability methods are often challenging to evaluate and compare. With a multitude of explainers available, practitioners must often compare and select explainers based on quan…