22 papers
New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
Shiyao Cui, QingLin Zhang, Di Wang +7
Neologisms, emerging terms in meaning or form, can serve as new vehicles for toxic expression, like "country girl" as a stigmatizing label targeting feminism. Such toxic neologisms…
Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine
Jiaxian Lv, Shiyao Cui, Yingkang Wang +3
Multi-image content has become an increasingly prevalent form of visual communication in social media, giving rise to a new safety issue, multi-image implicit toxicity (MIIT), wher…
RUBAS: Rubric-Based Reinforcement Learning for Agent Safety
Xian Qi Loye, Qinglin Su, Zhexin Zhang +5
The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text generation. Existing alignment…
SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation
Junxiao Yang, Minghao Zhang, Xiaoce Wang +4
Recent generative models can now produce visual artifacts with realistic embedded text and layouts, creating a new misinformation threat: synthetic credibility. We introduce SYNCRE…
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
Zhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang +8
Large Reasoning Models (LRMs) have achieved remarkable success on reasoning-intensive tasks such as mathematics and programming. However, their enhanced reasoning capabilities do n…
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
Zhexin Zhang, Yuhao Sun, Junxiao Yang +4
Fine-tuning on open-source Large Language Models (LLMs) with proprietary data is now a standard practice for downstream developers to obtain task-specific LLMs. Surprisingly, we re…