2 citations · 2 across the 9 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Mechanistic Origin of Moral Indifference in Language Models
Lingyu Li, Yan Teng, Yingchun Wang
Existing behavioral alignment techniques for Large Language Models (LLMs) often neglect the discrepancy between surface compliance and internal unaligned representations, leaving L…
cs.CL2025
LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models
Zhiyuan Ning, Tianle Gu, Jiaxin Song +8
The widespread adoption and increasing prominence of large language models (LLMs) in global technologies necessitate a rigorous focus on ensuring their safety across a diverse rang…
cs.CL2024★ 2 cited
MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts
Tianle Gu, Kexin Huang, Ruilin Luo +4
Large Language Models (LLMs) can memorize sensitive information, raising concerns about potential misuse. LLM Unlearning, a post-hoc approach to remove this information from traine…