10 papers
Capability-Routed Guard: Defending Large Reasoning Models Against Reasoning-Centric Jailbreaks
Yiyong Liu, Yixin Wu, Jun Sakuma
Large reasoning models (LRMs) expose a new safety failure mode: adversarial prompts can manipulate reasoning context, task decomposition, or capability interpretation so that harmf…
From Celebrities to Anyone: Characterizing AI Nudification Content, Technology, and Community Dynamics on 4chan
Chi Cui, Yixin Wu, Yang Zhang
AI nudification uses generative models to create synthetic non-consensual sexually explicit imagery (SNEACI) of real individuals. Prior work has examined dedicated nudification pla…
Peering Behind the Shield: Guardrail Identification in Large Language Models
Ziqing Yang, Yixin Wu, Rui Wen +2
With the rapid adoption of large language models (LLMs), conversational AI agents have become widely deployed across real-world applications. To enhance safety, these agents are of…
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
Xinyu Zhang, Yixin Wu, Boyang Zhang +4
Images shared on social media often expose geographic cues. While early geolocation methods required expert effort and lacked generalization, the rise of Large Vision Language Mode…
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
Yixin Wu, Rui Wen, Chi Cui +2
Inference attacks have been widely studied and offer a systematic risk assessment of ML services; however, their implementation and the attack parameters for optimal estimation are…
UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
Yiting Qu, Xinyue Shen, Yixin Wu +3
With the advent of text-to-image models and concerns about their misuse, developers are increasingly relying on image safety classifiers to moderate their generated unsafe images.…