2 papers
cs.CR2026
STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack
Xutao Mao, Liangjie Zhao, Tao Liu +3
Red-teaming Vision-Language Models is essential for identifying vulnerabilities where adversarial image-text inputs trigger toxic outputs. Existing approaches treat image generatio…
cs.CL2025
Detection, Classification, and Mitigation of Gender Bias in Large Language Models
Xiaoqing Cheng, Hongying Zan, Lulu Kong +2
With the rapid development of large language models (LLMs), they have significantly improved efficiency across a wide range of domains. However, recent studies have revealed that L…