papers

Publications (7)

cs.CL2026

They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It

Alex Kwon

When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished proje…

cs.CL2026

FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact

Alex Kwon

AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, w…

cs.CL2026

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak

Alex Kwon

The paper investigates why a short prefill phrase can disable the refusal behavior of aligned language models, pinpointing the failure to an early part of the model's response and…

#language model safety#jailbreak attacks#response‑site mechanisms#causal probing
cs.CL2026

Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression

Alex Kwon

Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. We ask wha…

cs.CL2026

Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One

Alex Kwon

A language model's memory can be worse than no memory at all when the model or its interface is disposed to act on it: a memory that keeps a wrong conclusion but drops the work beh…

cs.CR2026

Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts

Alex Kwon

LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation into stored "facts" that later steps tru…

cs.AI2026

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

Dasol Choi, Alex Kwon

Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational update flips which action i…