Publications (7)
They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It
Alex Kwon
When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished proje…
FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact
Alex Kwon
AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, w…
Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
Alex Kwon
The paper investigates why a short prefill phrase can disable the refusal behavior of aligned language models, pinpointing the failure to an early part of the model's response and…
Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
Alex Kwon
Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. We ask wha…
Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
Alex Kwon
A language model's memory can be worse than no memory at all when the model or its interface is disposed to act on it: a memory that keeps a wrong conclusion but drops the work beh…
Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts
Alex Kwon
LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation into stored "facts" that later steps tru…
When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models
Dasol Choi, Alex Kwon
Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational update flips which action i…