Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework
Hamidreza Saghir
Recent white-box OOD detection methods for LLMs -- including CED, RAUQ, and WildGuard confidence scores -- appear effective, but we show they are structurally confounded by sequenc…
cs.CL2025
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
Jingcheng Niu, Xingdi Yuan, Tong Wang +2
We observe a novel phenomenon, contextual entrainment, across a wide range of language models (LMs) and prompt settings, providing a new mechanistic perspective on how LMs become d…