3 papers
cs.SD2026
SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
Weilin Lin, Jianze Li, Hui Xiong +1
Large Audio-Language Models (LALMs) are becoming essential as a powerful multimodal backbone for real-world applications. However, recent studies show that audio inputs can more ea…
cs.CR2026
RedEdit: Agentic Red-Teaming of Image Safety Classifiers via MCTS-Guided Photo-Editing
Weilin Lin, Ziqi Lin, Zhenxing Zhou +4
Image safety classifiers serve as a critical component of contemporary content moderation systems on the internet. However, their resilience against user-style malicious image edit…
cs.CR2025
BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model
Weilin Lin, Nanjun Zhou, Yanyun Wang +3
Backdoor learning is a critical research topic for understanding the vulnerabilities of deep neural networks. While the diffusion model (DM) has been broadly deployed in public ove…