2 papers
cs.CR2026
Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs
Anirudh Sekar, Mrinal Agarwal, Rachel Sharma +4
Prompt injection attacks have become an increasing vulnerability for LLM applications, where adversarial prompts exploit indirect input channels such as emails or user-generated co…
cs.CL2025
TextBandit: Evaluating Probabilistic Reasoning in LLMs Through Language-Only Decision Tasks
Jimin Lim, Arjun Damerla, Arthur Jiang +1
Large language models (LLMs) have shown to be increasingly capable of performing reasoning tasks, but their ability to make sequential decisions under uncertainty only using natura…