3 papers
cs.AI2026
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Zhihao Zhang, Liting Huang, Guanghao Wu +3
Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal of benign queries or unsafe compl…
cs.CL2026
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
Shenyan Zheng, Jiayou Zhong, Anudeex Shetty +3
As large language models are increasingly used in high-stakes domains, it is essential that their outputs reflect not average} human preference, rather range of varying perspective…
cs.SI2025
Community Moderation and the New Epistemology of Fact Checking on Social Media
Isabelle Augenstein, Michiel Bakker, Tanmoy Chakraborty +13
Social media platforms have traditionally relied on internal moderation teams and partnerships with independent fact-checking organizations to identify and flag misleading content.…