Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Prompt Injection as Role Confusion
Charles Ye, Jasmine Cui, Dylan Hadfield-Menell
LLMs see the world as a single stream of text, partitioned into roles like <user> or <tool>. We trace prompt injection to role confusion: models perceive the source of text from ho…
cs.CL2026
Emergent Search and Backtracking in Latent Reasoning Models
Jasmine Cui, Charles Ye
What happens when a language model thinks without words? Standard reasoning LLMs verbalize intermediate steps as chain-of-thought; latent reasoning transformers (LRTs) instead perf…