5 papers
Robust Context-Aware Detection of Malicious Instructions in Text
Buzhao Liu, Xinhang Ma, Yevgeniy Vorobeychik
The remarkable instruction-following ability of modern LLMs has enabled their practical use as the minds of agents that can autonomously complete increasingly complex tasks. Therei…
AutoDojo: Adaptive Black-Box Attacks Reveal the Limits of IPI Defenses and Task-Specification Effects in LLM Agents
Xinhang Ma, Taoran Li, Chaowei Xiao +3
Indirect prompt injection (IPI) is a major security threat to LLM-powered agents. Thus, a growing body of work have proposed a variety of defensive approaches against IPI. These ca…
Protecting Language Models Against Unauthorized Distillation through Trace Rewriting
Xinhang Ma, William Yeoh, Ning Zhang +1
Knowledge distillation is a widely adopted technique for transferring capabilities from LLMs to smaller, more efficient student models. However, unauthorized use of knowledge disti…
Conformal Reachability for Safe Control in Unknown Environments
Xinhang Ma, Junlin Wu, Yiannis Kantaros +1
Designing provably safe control is a core problem in trustworthy autonomy. However, most prior work in this regard assumes either that the system dynamics are known or deterministi…
Learning Vision-Based Neural Network Controllers with Semi-Probabilistic Safety Guarantees
Xinhang Ma, Junlin Wu, Hussein Sibai +2
Ensuring safety in autonomous systems with vision-based control remains a critical challenge due to the high dimensionality of image inputs and the fact that the relationship betwe…