◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Alice Rigg

5 papers hereh-index 231 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2
  • last author3

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.CL1

identity via Semantic Scholar / OpenAlex

collaborators

5 papers

cs.LG2025

Distribution-Aware Feature Selection for SAEs

Narmeen Oozeer, Nirmalendu Prakash, Michael Lan +2

Sparse autoencoders (SAEs) decompose neural activations into interpretable features. A widely adopted variant, the TopK SAE, reconstructs each token from its K most active latents.…

cs.CL2025

Detecting and Characterizing Planning in Language Models

Jatin Nainani, Sankaran Vaidyanathan, Connor Watts +2

Modern large language models (LLMs) have demonstrated impressive performance across a wide range of multi-step reasoning tasks. Recent work suggests that LLMs may perform planning…

cs.LG2025

Bilinear MLPs enable weight-based mechanistic interpretability

Michael T. Pearce, Thomas Dooms, Alice Rigg +2

A mechanistic understanding of how MLPs do computation in deep neural networks remains elusive. Current interpretability work can extract features from hidden activations over an i…

cs.LG2025

Converting MLPs into Polynomials in Closed Form

Nora Belrose, Alice Rigg

Recent work has shown that purely quadratic functions can replace MLPs in transformers with no significant loss in performance, while enabling new methods of interpretability based…

cs.LG2024

Bilinear Convolution Decomposition for Causal RL Interpretability

Narmeen Oozeer, Sinem Erisken, Alice Rigg

Efforts to interpret reinforcement learning (RL) models often rely on high-level techniques such as attribution or probing, which provide only correlational insights and coarse cau…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.