#defense mechanisms

topicdefense mechanisms

6 papers · 1 filter

cs.CR2026

Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents

Mingxiao Liu, Yitong Li, Haoren Zhao +6

The paper studies stealthy audio prompt injection attacks that hide malicious instructions within normal speech to hijack multimodal LLM agents, introduces a benchmark (AudioAgentS…

cs.CR2026

Agent Harness Distillation: Inference-Time Harness Extraction and Exploitation in Autonomous Multi-Agent Systems

Yu Cui, Wuli Yang, Yirui Shi +4

The paper presents Agent Harness Distillation (AHD), a method for extracting and replicating inference-time coordination mechanisms (harnesses) from autonomous multi-agent systems…

cs.AI2026

MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck

Dongyi Liu, Haixing He, Xiaobao Wu +1

The paper introduces MIND, a lightweight framework that uses an intent‑aware information bottleneck to detect and filter poisoned memory in large language model agents, reducing at…

cs.CR2026

Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning

Hongliang Zhang, Zhongyuan Yu, Guijuan Wang +4

The paper proposes FedDAB, a two‑phase defense for federated learning that uses contrastive regularization and alignment checking to detect and exclude malicious local updates caus…

cs.CR2026

MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents

Jifeng Gao, Kang Xia, Yi Zhang +5

The paper introduces MemPoison, a benchmark and analysis framework that studies how adversarial content can be injected into the persistent external memory of large language model…

cs.CR2026

Context Contamination in LLM Analysis of Network Security Logs: Poison with Passive Prompt Injection and Mitigation Evaluation

Rabimba Karanjai, Yang Lu, Hemanth Hegadehalli Madhavarao +2

The paper shows that large language models used to analyze network security logs can be tricked by malicious log entries that inject hidden prompts, and it evaluates attacks and de…