22 papers
Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle
Jiaming Zhang, Boyang Chen, Zherui Li +14
Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Legal and regulatory remedies can address misuse, but many technica…
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
Zhouyuan Ma, Yutao Wu, Hanxun Huang +6
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is kn…
GeoDetect: Geometric Adversarial Detection for VLPs
Afsaneh Hasanebrahimi, Hanxun Huang, Christopher Leckie +2
The paper introduces GeoDetect, a method that uses geometric properties of vision‑language model embeddings to detect adversarial examples by measuring how far they deviate from th…
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
Xiang Zheng, Yutao Wu, Hanxun Huang +5
Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, and self-directed interaction. However, th…
Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs
Afsaneh Hasanebrahimi, Hanxun Huang, Christopher Leckie +1
Vision-Language models (VLMs), such as CLIP, achieve powerful zero-shot classification. However, their predictions remain sensitive to spurious correlations, where contextual cues…
When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries
Yige Li, Jun Sun, Wei Zhao +5
Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains poorly understood. We introduce…