#adversarial attacks
26 papers match
Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
Marco Alecci, Francesco Marchiori, Iyiola Emmanuel Olatunji +2
The paper evaluates three commercial image‑moderation services built on foundation models and shows that simple, model‑agnostic image transformations (e.g., color inversion, graysc…
Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
Fazhong Liu, Zhuoyan Chen, Haozhen Tan +3
The paper surveys security risks and defenses for world-model-based embodied AI, examining how attacks can affect data, perception, prediction, and action throughout the system’s l…
Revisiting the Adversarial Robustness of Graph-Based Traffic Forecasting
Qingzhao Zhang
The paper examines realistic, targeted adversarial attacks on graph-based traffic forecasting models and proposes a physics‑informed detection‑based defense that improves robustnes…
Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
Mingxiao Liu, Yitong Li, Haoren Zhao +6
The paper studies stealthy audio prompt injection attacks that hide malicious instructions within normal speech to hijack multimodal LLM agents, introduces a benchmark (AudioAgentS…
ToxScreen: Detecting Whether an LLM Has Been Poisoned
Anthony Hughes, Nicole Xing, Collin Francel +2
The paper introduces ToxScreen, a benchmark of backdoored large language models, and evaluates methods for recovering hidden triggers under realistic defender constraints, finding…
Pangram 4 Technical Report
Ben Glickenhaus, Katherine Thai, Jenna Russell +4
The paper introduces Pangram 4, a deep‑learning model for detecting AI‑generated text that achieves high accuracy, strong out‑of‑distribution robustness, and improved detection of…
Physically Real-time Infrared Attack against Optical Flow Estimation Networks
Shen You, Wei Jiang, Jiarui Liu +4
The paper proposes a real-time physical attack using infrared lights to generate adversarial examples that fool optical flow estimation networks without modifying the victim system…
IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks
Mengqi He, Jing Zhang
The paper proposes IGME, an efficient method that chains attack components to generate transferable adversarial perturbations for semantic segmentation using only a single source m…
Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models
Saurabh Yadav, Badri Narayana Patro, Vijay Srinivas Agneeswaran
The paper evaluates how diffusion-based language models handle noisy inputs and adversarial attacks compared to traditional autoregressive models, finding that while they resist ce…
Evaluation of Adversarial Robustness in Arabic Language Models
Anwar Alajmi, Ayed Salman, Imtiaz Ahmad
The paper evaluates how vulnerable five Arabic language models are to various adversarial attacks at character, word, and sentence levels, and examines how adversarial training can…
I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models
Yimao Guo, Zuomin Qu, Wei Lu
The paper introduces I2VShield, a lightweight proactive defense that generates text‑adaptive perturbations and uses a multimodal attention disruption attack to protect images from…
CluCERT: Certifying LLM Robustness via Clustering-Guided Denoising Smoothing
Zixia Wang, Gaojie Jin, Jia Hu +1
The paper presents CluCERT, a framework that uses clustering-guided denoising smoothing to certify the robustness of large language models against adversarial synonym substitutions…
ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho +3
ARMOR++ is a multi‑agent system that uses vision‑language and large language models to coordinate several attack primitives, creating highly transferable adversarial examples that…
Pretraining Data Can Be Poisoned through Computational Propaganda
Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith +2
The paper shows that language model pretraining data can be poisoned through publicly editable web discussion pages, and introduces a method called HalfLife to estimate how much ma…
Random Logit Scaling: Defending Deep Neural Networks Against Black-Box Score-Based Adversarial Example Attacks
Hamid Dashtbani, Mehdi Dousti Gandomani, AmirMahdi Sadeghzadeh
The paper introduces Random Logit Scaling, a plug‑and‑play post‑processing defense that randomly rescales model logits to thwart black‑box score‑based adversarial attacks while kee…
On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline
Yuchen Ren, Zhengyu Zhao, Chenhao Lin +2
The paper introduces SimVLA, a simplified vision‑language adversarial attack pipeline that improves transferability and computational efficiency compared to existing complex method…
Power Homotopy for Zeroth-Order Non-Convex Optimizations
Chen Xu
The paper introduces GS-PowerHP, a homotopy-based method that gradually reduces the Gaussian smoothing radius to improve exploration and refinement in zeroth-order non-convex optim…
Evaluating Frontier AI Agents as Autonomous Clinical Security Auditors
Michael O. Eniolade
The paper introduces an evaluation framework where frontier AI agents autonomously perform security audits on clinical prediction models by executing adversarial attacks, computing…
UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors
Dima Galat, Marian-Andrei Rizoiu
The paper examines how AI-generated text can evade state-of-the-art detection models by deliberately moving the output outside the detectors' training distribution, introducing two…
Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection
Zhenpeng Li
The paper introduces Traffic-Aware Randomized Smoothing, a method that adds Gaussian noise only to attacker‑controllable network traffic features during fine‑tuning and certificati…
Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation
Mohammad Allahbakhsh, Mohammad Hassan Bahari, Moslem Attar-Raouf
The paper proposes a new way to conduct penetration testing for AI-enabled systems by focusing on inducing undesirable AI-driven behavior that violates operational objectives, rath…
Adversarial Attacks on Online Handwriting using Salience-based Temporal Editing
Yataro Tamura, Brian Kenji Iwana, Jiseok Lee
The paper introduces a new adversarial attack for online handwriting recognition that edits the pen trajectory by inserting or deleting points guided by temporal salience, preservi…
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
Chenyang Li, Kaige Li, Zeyu Jiang +1
The paper introduces AdvNav, a gradient‑free black‑box adversarial attack that perturbs first‑person visual inputs to disrupt vision‑and‑language navigation agents, using behavior‑…
When cheap gradients fail: the measurement cost of attacking quantum classifiers
Bacui Li, Chandra Thapa, Tansu Alpcan +1
The paper shows that shot noise from finite quantum measurements creates a natural defense against gradient-based adversarial attacks on variational quantum classifiers, requiring…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.