2 citations · 5 across the 16 of their papers we have counts for
18 papers
Speculative Probing: LLM Monitoring at Speculative-Decoding Cost
Collin Zhang, Tingwei Zhang, Vitaly Shmatikov
Real-time classification during language model inference is valuable for safety filtering, behavioral analysis, and model monitoring, but current approaches force a trade-off betwe…
Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
Rishi Jha, Harold Triedman, Arkaprabha Bhattacharya +1
Agents operating with computer and Web use inevitably encounter errors: inaccessible webpages, missing files, local and remote misconfigurations, etc. These errors do not thwart ag…
Deep-Research Agents Can Be Poisoned via User-Generated Content
Tingwei Zhang, Harold Triedman, Vitaly Shmatikov
Deep-research agents are an alternative to conventional Web search. They use multi-agent pipelines to issue multiple Web searches related to user queries, retrieve relevant content…
How to Steal Reasoning Without Reasoning Traces
Tingwei Zhang, John X. Morris, Vitaly Shmatikov
Many large language models (LLMs) use reasoning to generate responses but do not reveal their full reasoning traces (a.k.a. chains of thought), instead outputting only final answer…
Learning to Detect Language Model Training Data via Active Reconstruction
Junjie Oscar Yin, John X. Morris, Vitaly Shmatikov +2
Detecting LLM training data is generally framed as a membership inference attack (MIA) problem. However, conventional MIAs operate passively on fixed model weights, using log-likel…
Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems
Rishi Jha, Harold Triedman, Justin Wagle +1
Control-flow hijacking attacks manipulate orchestration mechanisms in multi-agent systems into performing unsafe actions that compromise the system and exfiltrate sensitive informa…