3 papers
cs.CV2026
Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks
Yunqi Xue, Zhijiang Li, Philip Torr +1
Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discretized visual tokens…
cs.CR2026
LLM Jailbreak Detection for (Almost) Free!
Guorui Chen, Yifan Xia, Xiaojun Jia +3
Large language models (LLMs) enhance security through alignment when widely used, but remain susceptible to jailbreak attacks capable of producing inappropriate content. Jailbreak…
cs.AI2025
Reimagining Safety Alignment with An Image
Yifan Xia, Guorui Chen, Wenqian Yu +3
Large language models (LLMs) excel in diverse applications but face dual challenges: generating harmful content under jailbreak attacks and over-refusal of benign queries due to ri…