◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Eoin Farrell

2 papers hereh-index 2151 citations2 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1

Across the 1 of 2 papers where every author was matched, so the position is known.

fields
  • cs.LG2

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.LGShow all

2 papers · 1 filter

cs.LG2025

SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Adam Karvonen, Can Rager, Johnny Lin +12

Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most pri…

cs.LG2024

Applying sparse autoencoders to unlearn knowledge in language models

Eoin Farrell, Yeu-Tong Lau, Arthur Conmy

We investigate whether sparse autoencoders (SAEs) can be used to remove knowledge from language models. We use the biology subset of the Weapons of Mass Destruction Proxy dataset a…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.