7 papers
Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
Alex Kwon
Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. We ask wha…
FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact
Alex Kwon
AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, w…
Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
Alex Kwon
A language model's memory can be worse than no memory at all when the model or its interface is disposed to act on it: a memory that keeps a wrong conclusion but drops the work beh…
Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
Alex Kwon
Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal. We ask where and how it fails. The harm representation stays intact: o…
They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It
Alex Kwon
When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished proje…
Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts
Alex Kwon
LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation into stored "facts" that later steps tru…