Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
Rui Li, Peiyi Wang, Jingyuan Ma +3
Large Language Models (LLMs) have gained increasing attention for their remarkable capacity, alongside concerns about safety arising from their potential to produce harmful content…
cs.CL2024
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
Zhexin Zhang, Yida Lu, Jingyuan Ma +8
The safety of Large Language Models (LLMs) has gained increasing attention in recent years, but there still lacks a comprehensive approach for detecting safety issues within LLMs'…