Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment
Anurag Kumar, Raghuveer Peri, Jon Burnsky +4
Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Models aligned on text alone show a highe…
cs.LG2025
A Closer Look at Adversarial Suffix Learning for Jailbreaking LLMs: Augmented Adversarial Trigger Learning
Zhe Wang, Yanjun Qi
Gradient optimization-based adversarial attack methods automate the learning of adversarial triggers to generate jailbreak prompts or leak system prompts. In this work, we take a c…
cs.LG2024
Hierarchical Prompt Decision Transformer: Improving Few-Shot Policy Generalization with Global and Adaptive Guidance
Zhe Wang, Haozhu Wang, Yanjun Qi
Decision transformers recast reinforcement learning as a conditional sequence generation problem, offering a simple but effective alternative to traditional value or policy-based m…