3 citations · 3 across the 1 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Contrastive Weak-to-strong Generalization
Houcheng Jiang, Junfeng Fang, Jiaxin Wu +5
Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples from aligned weaker ones, without requir…
cs.CL2025
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
Houcheng Jiang, Zetong Zhao, Junfeng Fang +5
Safety-aligned large language models (LLMs) remain vulnerable to backdoor attacks. Recent model editing-based approaches enable efficient backdoor injection by directly modifying a…
cs.CL2023★ 3 cited
Attack Prompt Generation for Red Teaming and Defending Large Language Models
Boyi Deng, Wenjie Wang, Fuli Feng +3
Large language models (LLMs) are susceptible to red teaming attacks, which can induce LLMs to generate harmful content. Previous research constructs attack prompts via manual or au…