activity
20242026
most citedMachine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents

2 citations · 5 across the 16 of their papers we have counts for

collaborators

18 papers

cs.AI2026

Speculative Probing: LLM Monitoring at Speculative-Decoding Cost

Collin Zhang, Tingwei Zhang, Vitaly Shmatikov

Real-time classification during language model inference is valuable for safety filtering, behavioral analysis, and model monitoring, but current approaches force a trade-off betwe…

cs.CL2026

Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents

Rishi Jha, Harold Triedman, Arkaprabha Bhattacharya +1

Agents operating with computer and Web use inevitably encounter errors: inaccessible webpages, missing files, local and remote misconfigurations, etc. These errors do not thwart ag…

cs.CR2026

Deep-Research Agents Can Be Poisoned via User-Generated Content

Tingwei Zhang, Harold Triedman, Vitaly Shmatikov

Deep-research agents are an alternative to conventional Web search. They use multi-agent pipelines to issue multiple Web searches related to user queries, retrieve relevant content…

cs.CR2026

How to Steal Reasoning Without Reasoning Traces

Tingwei Zhang, John X. Morris, Vitaly Shmatikov

Many large language models (LLMs) use reasoning to generate responses but do not reveal their full reasoning traces (a.k.a. chains of thought), instead outputting only final answer…

cs.LG2026

Learning to Detect Language Model Training Data via Active Reconstruction

Junjie Oscar Yin, John X. Morris, Vitaly Shmatikov +2

Detecting LLM training data is generally framed as a membership inference attack (MIA) problem. However, conventional MIAs operate passively on fixed model weights, using log-likel…

cs.LG2025

Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems

Rishi Jha, Harold Triedman, Justin Wagle +1

Control-flow hijacking attacks manipulate orchestration mechanisms in multi-agent systems into performing unsafe actions that compromise the system and exfiltrate sensitive informa…