◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Dan Mossing

2 papers hereh-index 266 citations2 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • last author1

Across the 1 of 2 papers where every author was matched, so the position is known.

fields
  • cs.LG2

identity via Semantic Scholar / OpenAlex

collaborators

2 papers

cs.LG2025

Persona Features Control Emergent Misalignment

Miles Wang, Tom Dupré la Tour, Olivia Watkins +8

Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that…

cs.LG2025

Investigating task-specific prompts and sparse autoencoders for activation monitoring

Henk Tillman, Dan Mossing

Language models can behave in unexpected and unsafe ways, and so it is valuable to monitor their outputs. Internal activations of language models encode additional information that…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.