Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Measurement of LLM's Philosophies of Human Nature
Minheng Ni, Ennan Wu, Zidong Gong +6
The widespread application of artificial intelligence (AI) in various tasks, along with frequent reports of conflicts or violations involving AI, has sparked societal concerns abou…
cs.CL2024
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Lijun Li, Bowen Dong, Ruohui Wang +5
In the rapidly evolving landscape of Large Language Models (LLMs), ensuring robust safety measures is paramount. To meet this crucial need, we propose \emph{SALAD-Bench}, a safety…