3 papers
cs.LG2026
Mechanisms of Introspective Awareness
Uzay Macar, Li Yang, Atticus Wang +3
Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introsp…
cs.CR2026
Invisible Hands: Gray-Box Bit Flip Attack for Steering LLMs Without Knowledge of Gradients, Data, and Weights
Abeer Matar A. Almalky, Ziyan Wang, Mohaiminul Al Nahian +2
In recent years, large language models (LLMs) have achieved remarkable advances and are increasingly deployed in critical applications across diverse domains. This growing adoption…
cs.CR2026
CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs
Mohaiminul Al Nahian, Abeer Matar A. Almalky, Gamana Aragonda +6
The rapid advancement of large language models (LLMs) has sparked growing interest in understanding their security vulnerabilities, particularly Trojan attacks that enable stealthy…