◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

David Krueger

8 papers hereh-index 8740 citations11 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author7

Across the 7 of 8 papers where every author was matched, so the position is known.

fields
  • cs.LG5
  • cs.CR1
  • cs.CV1
  • cs.CY1
same name
  • David Krueger — 16 papers, h 24
  • David Krueger — 8 papers, h 4
  • David Krueger — 6 papers
  • David Krueger — 6 papers, h 6
  • David Krueger — 5 papers, h 3
  • David Krueger — 5 papers, h 2

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20232025
most citedEnhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders

4 citations · 4 across the 3 of their papers we have counts for

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2025

Rethinking Safety in LLM Fine-tuning: An Optimization Perspective

Minseon Kim, Jin Myung Kwak, Lama Alssum +5

Fine-tuning language models is commonly believed to inevitably harm their safety, i.e., refusing to respond to harmful user requests, even when using harmless datasets, thus requir…

cs.LG2025

Open Problems in Machine Unlearning for AI Safety

Fazl Barez, Tingchen Fu, Ameya Prabhu +16

As AI systems become more capable, widely deployed, and increasingly autonomous in critical areas such as cybersecurity, biological research, and healthcare, ensuring their safety…

cs.LG2024★ 4 cited

Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders

Luke Marks, Alasdair Paren, David Krueger +1

Sparse Autoencoders (SAEs) have shown promise in improving the interpretability of neural network activations, but can learn features that are not features of the input, limiting t…

cs.LG2024

Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders

Michael Lan, Philip Torr, Austin Meek +3

The Universality Hypothesis in large language models (LLMs) claims that different models converge towards similar concept representations in their latent spaces. Providing evidence…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.