◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Todd Nief

4 papers hereh-index 318 citations9 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author3

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.AI1
  • cs.CL1

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators

4 papers

cs.AI2026

Iterative Finetuning is Mostly Idempotent

Zephaniah Roe, Jack Sanderson, Dang Nguyen +5

If a model has some behavioral tendency, such as sycophancy or misalignment, and it is trained on its own outputs, will the tendency be amplified in the next generation of models?…

cs.LG2025

Dynamic Weight Grafting: Localizing Finetuned Factual Knowledge in Transformers

Todd Nief, David Reber, Sean Richardson +1

When an LLM learns a new fact during finetuning (e.g., new movie releases, newly elected pope, etc.), where does this information go? Are entities enriched with relation informatio…

cs.LG2024

Deep Model Merging: The Sister of Neural Network Interpretability -- A Survey

Arham Khan, Todd Nief, Nathaniel Hudson +6

We survey the model merging literature through the lens of loss landscape geometry to connect observations from empirical studies on model merging and loss landscape analysis to ph…

cs.CL2024

RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals

David Reber, Sean Richardson, Todd Nief +2

Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.