5 papers
Speculative Probing: LLM Monitoring at Speculative-Decoding Cost
Collin Zhang, Tingwei Zhang, Vitaly Shmatikov
Real-time classification during language model inference is valuable for safety filtering, behavioral analysis, and model monitoring, but current approaches force a trade-off betwe…
How to Steal Reasoning Without Reasoning Traces
Tingwei Zhang, John X. Morris, Vitaly Shmatikov
Many large language models (LLMs) use reasoning to generate responses but do not reveal their full reasoning traces (a.k.a. chains of thought), instead outputting only final answer…
Adversarial Hubness in Multi-Modal Retrieval
Tingwei Zhang, Fnu Suya, Rishi Jha +2
Hubness is a phenomenon in high-dimensional vector spaces where a point from the natural distribution is unusually close to many other points. This is a well-known problem in infor…
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives
Collin Zhang, Tingwei Zhang, Vitaly Shmatikov
We design, implement, and evaluate adversarial decoding, a new, generic text generation technique that produces readable documents for different adversarial objectives. Prior metho…
Self-interpreting Adversarial Images
Tingwei Zhang, Collin Zhang, John X. Morris +2
We introduce a new type of indirect, cross-modal injection attacks against visual language models that enable creation of self-interpreting images. These images contain hidden "met…