11 citations · 17 across the 28 of their papers we have counts for
Showing cs.CRShow all
2 papers · 1 filter
cs.CR2026
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
Junbo Zhang, Qianli Zhou, Xinyang Deng +3
Large language models (LLMs) suffer from degraded safety capabilities even when fine-tuned with benign datasets. However, existing methods for identifying safety-degrading samples…
cs.CR2025
Understanding and Mitigating Over-refusal for Large Language Models via Representation Intervention
Junbo Zhang, Ran Chen, Qianli Zhou +2
Large language models (LLMs) demonstrate powerful capabilities across various natural language processing tasks,yet their inherent safety vulnerabilities undermine the reliable app…