◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Adam Karvonen

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author1
  • first author2

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG3

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2025

Robustly Improving LLM Fairness in Realistic Settings via Interpretability

Adam Karvonen, Samuel Marks

Large language models (LLMs) are increasingly deployed in high-stakes hiring applications, making decisions that directly impact people's careers and livelihoods. While prior studi…

cs.LG2025

Revisiting End-To-End Sparse Autoencoder Training: A Short Finetune Is All You Need

Adam Karvonen

Sparse autoencoders (SAEs) are widely used for interpreting language model activations. A key evaluation metric is the increase in cross-entropy loss between the original model log…

cs.LG2024

Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks

Adam Karvonen, Can Rager, Samuel Marks +1

Sparse Autoencoders (SAEs) are an interpretability technique aimed at decomposing neural network activations into interpretable units. However, a major bottleneck for SAE developme…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.