◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Max Torop

5 papers hereh-index 326 citations9 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3
  • middle author2

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.CV1

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2026

Inverted Detection and Control in Steering Vectors

Max Torop, Aria Masoomi, Jennifer Dy

Steering vectors (SVs) are widely used to influence the expression of concepts (e.g., truthfulness) in large language model outputs. A key assumption underpinning SVs is that they…

cs.LG2025

DISCO: Disentangled Communication Steering for Large Language Models

Max Torop, Aria Masoomi, Masih Eskandar +1

A variety of recent methods guide large language model outputs via the inference-time addition of steering vectors to residual-stream or attention-head representations. In contrast…

cs.LG2025

Axiomatic Explainer Globalness via Optimal Transport

Davin Hill, Josh Bone, Aria Masoomi +2

Explainability methods are often challenging to evaluate and compare. With a multitude of explainers available, practitioners must often compare and select explainers based on quan…

cs.LG2024

Boundary-Aware Uncertainty for Feature Attribution Explainers

Davin Hill, Aria Masoomi, Max Torop +2

Post-hoc explanation methods have become a critical tool for understanding black-box classifiers in high-stakes applications. However, high-performing classifiers are often highly…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.