#hidden state analysis

try —

4 papers match

cs.AI2026

Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures

SiYuan Ma, Yiqin Luo, Zhangji +8

The paper introduces a technique called forked futures that samples future operations after a prefix state to compare hidden states of language models, enabling the discovery of re…

#language models#causal inference#hidden state analysis#model interpretability
cs.CL2026

Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States

Weiyi Kong, Zhuoran Li

The paper investigates whether transformer language models encode information about how a piece of diagnostic evidence supports, challenges, or is unrelated to a causal claim, usin…

#causal inference#language model probing#prompt engineering#hidden state analysis
cs.LG2026

Code Correctness Is Linearly Decodable from LLM Hidden States Before Generation

Carlo Di Cicco

The paper shows that the hidden state of a large language model right before it starts generating code contains a linear signal that predicts whether the produced code will be corr…

#code generation#large language models#hidden state analysis#linear probing
cs.CR2026

PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis

Junhui Wang, Hangtao Zhang, Zhirun Zheng +5

The paper introduces PVDetector, a training‑free method that detects prompt injection attacks on purpose‑specific LLM agents by measuring alignment of hidden states with policy‑vio…

#prompt injection#large language models#policy violation detection#agent security

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.