◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Nora Belrose

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • last author2

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG3
same name
  • Nora Belrose — 5 papers
  • Nora Belrose — 1 paper, h 3

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedDoes Transformer Interpretability Transfer to RNNs?

2 citations · 2 across the 3 of their papers we have counts for

collaborators

3 papers

cs.LG2024

Understanding Gradient Descent through the Training Jacobian

Nora Belrose, Adam Scherlis

We examine the geometry of neural network training using the Jacobian of trained network parameters with respect to their initial values. Our analysis reveals low-dimensional struc…

cs.LG2024

Refusal in LLMs is an Affine Function

Thomas Marshall, Adam Scherlis, Nora Belrose

We propose affine concept editing (ACE) as an approach for steering language models' behavior by intervening directly in activations. We begin with an affine decomposition of model…

cs.LG2024★ 2 cited

Does Transformer Interpretability Transfer to RNNs?

Gonçalo Paulo, Thomas Marshall, Nora Belrose

Recent advances in recurrent neural network architectures, such as Mamba and RWKV, have enabled RNNs to match or exceed the performance of equal-size transformers in terms of langu…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.