◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Andrew Mack

3 papers hereh-index 1259 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • hep-th1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2026

Scaling Interpretable Transformers with Parity Bottleneck Layers

Andrew Mack, Kraig Yuheng Tou, Mark Henry +2

Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are de…

cs.LG2026

Mechanistically Eliciting Latent Behaviors in Language Models

Andrew Mack, Nina Panickssery, Alexander Matt Turner

We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes. Such perturbations could help reshape model behavior and systemat…

hep-th2026

Towards Worst-Case Guarantees with Scale-Aware Interpretability

Lauren Greenspan, David Berman, Aryeh Brill +9

Neural networks organize information according to the hierarchical, multi-scale structure of natural data. Methods to interpret model internals should be similarly scale-aware, exp…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.