2 papers
cs.CL2026
When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints
Yuheng Chen, Zhiyu Wu, Bowen Cheng +1
Safety alignment in large language models (LLMs) is primarily evaluated under open-ended generation, where models can mitigate risk by refusing to respond. In contrast, many real-w…
cs.CL2025
AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output
Hisami Suzuki, Satoru Katsumata, Takashi Kodama +3
In this paper we present AnswerCarefully, a dataset for promoting the safety and appropriateness of Japanese LLM outputs. The dataset consists of 1,800 pairs of questions and refer…