Showing cs.CRShow all
2 papers · 1 filter
cs.CR2026
DREAM: Dynamic Red-teaming across Environments for AI Models
Liming Lu, Xiang Gu, Junyu Huang +5
Large Language Models (LLMs) are increasingly used in agentic systems, where their interactions with diverse tools and environments create complex, multi-stage safety challenges. H…
cs.CR2025
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
Xiyu Zeng, Siyuan Liang, Liming Lu +5
As the capabilities of Vision Language Models (VLMs) continue to improve, they are increasingly targeted by jailbreak attacks. Existing defense methods face two major limitations:…