#adversarial attacks

try —

26 papers match

cs.AI2026

Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation

Marco Alecci, Francesco Marchiori, Iyiola Emmanuel Olatunji +2

The paper evaluates three commercial image‑moderation services built on foundation models and shows that simple, model‑agnostic image transformations (e.g., color inversion, graysc…

#image moderation#adversarial attacks#foundation models#robustness evaluation
cs.CR2026

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

Fazhong Liu, Zhuoyan Chen, Haozhen Tan +3

The paper surveys security risks and defenses for world-model-based embodied AI, examining how attacks can affect data, perception, prediction, and action throughout the system’s l…

#embodied ai#world models#adversarial attacks#defense strategies
cs.CR2026

Revisiting the Adversarial Robustness of Graph-Based Traffic Forecasting

Qingzhao Zhang

The paper examines realistic, targeted adversarial attacks on graph-based traffic forecasting models and proposes a physics‑informed detection‑based defense that improves robustnes…

#graph neural networks#traffic forecasting#adversarial attacks#physics-informed detection
cs.CR2026

Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents

Mingxiao Liu, Yitong Li, Haoren Zhao +6

The paper studies stealthy audio prompt injection attacks that hide malicious instructions within normal speech to hijack multimodal LLM agents, introduces a benchmark (AudioAgentS…

#audio injection#multimodal llm#prompt injection#adversarial attacks
cs.CR2026

ToxScreen: Detecting Whether an LLM Has Been Poisoned

Anthony Hughes, Nicole Xing, Collin Francel +2

The paper introduces ToxScreen, a benchmark of backdoored large language models, and evaluates methods for recovering hidden triggers under realistic defender constraints, finding…

#large language models#backdoor detection#adversarial attacks#model security
cs.CL2026

Pangram 4 Technical Report

Ben Glickenhaus, Katherine Thai, Jenna Russell +4

The paper introduces Pangram 4, a deep‑learning model for detecting AI‑generated text that achieves high accuracy, strong out‑of‑distribution robustness, and improved detection of…

#ai text detection#deep learning#model robustness#adversarial attacks
cs.CV2026

Physically Real-time Infrared Attack against Optical Flow Estimation Networks

Shen You, Wei Jiang, Jiarui Liu +4

The paper proposes a real-time physical attack using infrared lights to generate adversarial examples that fool optical flow estimation networks without modifying the victim system…

#adversarial attacks#optical flow#infrared illumination#real-time
cs.CV2026

IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks

Mengqi He, Jing Zhang

The paper proposes IGME, an efficient method that chains attack components to generate transferable adversarial perturbations for semantic segmentation using only a single source m…

#adversarial attacks#semantic segmentation#transferability#ensemble methods
cs.AI2026

Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models

Saurabh Yadav, Badri Narayana Patro, Vijay Srinivas Agneeswaran

The paper evaluates how diffusion-based language models handle noisy inputs and adversarial attacks compared to traditional autoregressive models, finding that while they resist ce…

#diffusion language models#robustness#adversarial attacks#calibration
cs.CL2026

Evaluation of Adversarial Robustness in Arabic Language Models

Anwar Alajmi, Ayed Salman, Imtiaz Ahmad

The paper evaluates how vulnerable five Arabic language models are to various adversarial attacks at character, word, and sentence levels, and examines how adversarial training can…

#adversarial attacks#arabic language models#model robustness#defense techniques
cs.CV2026

I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models

Yimao Guo, Zuomin Qu, Wei Lu

The paper introduces I2VShield, a lightweight proactive defense that generates text‑adaptive perturbations and uses a multimodal attention disruption attack to protect images from…

#image-to-video generation#proactive defense#adversarial attacks#diffusion transformer
cs.LG2026

CluCERT: Certifying LLM Robustness via Clustering-Guided Denoising Smoothing

Zixia Wang, Gaojie Jin, Jia Hu +1

The paper presents CluCERT, a framework that uses clustering-guided denoising smoothing to certify the robustness of large language models against adversarial synonym substitutions…

#large language models#robustness certification#adversarial attacks#denoising smoothing
cs.CV2026

ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors

Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho +3

ARMOR++ is a multi‑agent system that uses vision‑language and large language models to coordinate several attack primitives, creating highly transferable adversarial examples that…

#adversarial attacks#deepfake detection#transferability#multi-agent orchestration
cs.AI2026

Pretraining Data Can Be Poisoned through Computational Propaganda

Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith +2

The paper shows that language model pretraining data can be poisoned through publicly editable web discussion pages, and introduces a method called HalfLife to estimate how much ma…

#data poisoning#language model pretraining#web crawling#adversarial attacks
cs.LG2026

Random Logit Scaling: Defending Deep Neural Networks Against Black-Box Score-Based Adversarial Example Attacks

Hamid Dashtbani, Mehdi Dousti Gandomani, AmirMahdi Sadeghzadeh

The paper introduces Random Logit Scaling, a plug‑and‑play post‑processing defense that randomly rescales model logits to thwart black‑box score‑based adversarial attacks while kee…

#adversarial attacks#black-box defense#logit scaling#randomization
cs.CV2026

On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline

Yuchen Ren, Zhengyu Zhao, Chenhao Lin +2

The paper introduces SimVLA, a simplified vision‑language adversarial attack pipeline that improves transferability and computational efficiency compared to existing complex method…

#adversarial attacks#vision-language models#transferability#efficiency
math.OC2026

Power Homotopy for Zeroth-Order Non-Convex Optimizations

Chen Xu

The paper introduces GS-PowerHP, a homotopy-based method that gradually reduces the Gaussian smoothing radius to improve exploration and refinement in zeroth-order non-convex optim…

#zeroth-order optimization#non-convex optimization#smoothing#homotopy
cs.CR2026

Evaluating Frontier AI Agents as Autonomous Clinical Security Auditors

Michael O. Eniolade

The paper introduces an evaluation framework where frontier AI agents autonomously perform security audits on clinical prediction models by executing adversarial attacks, computing…

#clinical AI auditing#adversarial attacks#automated security evaluation#large language models
cs.CR2026

UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors

Dima Galat, Marian-Andrei Rizoiu

The paper examines how AI-generated text can evade state-of-the-art detection models by deliberately moving the output outside the detectors' training distribution, introducing two…

#ai text detection#adversarial attacks#out-of-distribution generation#language model evasion
cs.CR2026

Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection

Zhenpeng Li

The paper introduces Traffic-Aware Randomized Smoothing, a method that adds Gaussian noise only to attacker‑controllable network traffic features during fine‑tuning and certificati…

#intrusion detection#large language models#randomized smoothing#certified robustness
cs.CR2026

Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation

Mohammad Allahbakhsh, Mohammad Hassan Bahari, Moslem Attar-Raouf

The paper proposes a new way to conduct penetration testing for AI-enabled systems by focusing on inducing undesirable AI-driven behavior that violates operational objectives, rath…

#penetration testing#ai security#adversarial attacks#prompt injection
cs.LG2026

Adversarial Attacks on Online Handwriting using Salience-based Temporal Editing

Yataro Tamura, Brian Kenji Iwana, Jiseok Lee

The paper introduces a new adversarial attack for online handwriting recognition that edits the pen trajectory by inserting or deleting points guided by temporal salience, preservi…

#adversarial attacks#online handwriting#temporal editing#salience mapping
cs.AI2026

AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation

Chenyang Li, Kaige Li, Zeyu Jiang +1

The paper introduces AdvNav, a gradient‑free black‑box adversarial attack that perturbs first‑person visual inputs to disrupt vision‑and‑language navigation agents, using behavior‑…

#vision-language navigation#adversarial attacks#black-box optimization#embodied AI
quant-ph2026

When cheap gradients fail: the measurement cost of attacking quantum classifiers

Bacui Li, Chandra Thapa, Tansu Alpcan +1

The paper shows that shot noise from finite quantum measurements creates a natural defense against gradient-based adversarial attacks on variational quantum classifiers, requiring…

#adversarial attacks#quantum classifiers#gradient estimation#measurement cost

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.