◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Sam Marks

2 papers

No researched profile yet.

papers

Publications (2)

cs.AI2026

Introspection Adapters: Training LLMs to Report Their Learned Behaviors

Keshav Shenoy, Li Yang, Abhay Sheshadri +4

When model developers or users fine-tune an LLM, this can induce behaviors that are unexpected, deliberately harmful, or hard to detect. It would be far easier to audit LLMs if the…

cs.AI2024

Alignment faking in large language models

Ryan Greenblatt, Carson Denison, Benjamin Wright +17

We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Sign in
  • Library
  • Chat
Data
  • arXiv.org
  • Latest RSS
Not affiliated with arXiv