3 papers
cs.CL2025
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
Huizhen Shu, Xuying Li, Qirui Wang +3
With the rapid proliferation of Natural Language Processing (NLP), especially Large Language Models (LLMs), generating adversarial examples to jailbreak LLMs remains a key challeng…
cs.CL2025
Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach
Xuying Li, Zhuo Li, Yuji Kosuga +1
Aligning large language models (LLMs) with human values and safety constraints is challenging, especially when objectives like helpfulness, truthfulness, and avoidance of harm conf…
cs.CL2025
Output Length Effect on DeepSeek-R1's Safety in Forced Thinking
Xuying Li, Zhuo Li, Yuji Kosuga +1
Large Language Models (LLMs) have demonstrated strong reasoning capabilities, but their safety under adversarial conditions remains a challenge. This study examines the impact of o…