2 papers
cs.CR2026
RLSpoofer: A Lightweight Evaluator for LLM Watermark Spoofing Resilience
Hanbo Huang, Xuan Gong, Yiran Zhang +2
Large language model (LLM) watermarking has emerged as a promising approach for detecting and attributing AI-generated text, yet its robustness to black-box spoofing remains insuff…
cs.CV2026
JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization
Haolun Zheng, Yu He, Tailun Chen +6
Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed…