Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
Yan Wang, Zhixuan Chu, Zihao Xue +9
Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do not always lead to faithful…
cs.CL2025
Mitigating Social Bias in Large Language Models: A Multi-Objective Approach within a Multi-Agent Framework
Zhenjie Xu, Wenqing Chen, Yi Tang +6
Natural language processing (NLP) has seen remarkable advancements with the development of large language models (LLMs). Despite these advancements, LLMs often produce socially bia…