activity
20232026
most citedShadow Alignment: The Ease of Subverting Safely-Aligned Language Models

10 citations · 11 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI2026

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

Yinhao Tang, Youqing Fang, Yanan Sun +8

Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reasoning-rich content but lack exp…

cs.HC2026

MindCopilot: Towards Formalizing and Evaluating Granular Human-LLM Co-Writing

Youqing Fang, Yinhao Tang, Yanan Sun +8

Recent writing assistants are increasingly shifting from passive, prompt-driven interaction to proactive, suggestion-based completion, which integrates localized continuations into…

cs.CL20241 cited

Unveiling the Misuse Potential of Base Large Language Models via In-Context Learning

Xiao Wang, Tianze Chen, Xianjun Yang +3

The open-sourcing of large language models (LLMs) accelerates application development, innovation, and scientific progress. This includes both base models, which are pre-trained on…

cs.CL2024

Navigating the OverKill in Large Language Models

Chenyu Shi, Xiao Wang, Qiming Ge +7

Large language models are meticulously aligned to be both helpful and harmless. However, recent research points to a potential overkill which means models may refuse to answer beni…

cs.CL202310 cited

Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Xianjun Yang, Xiao Wang, Qi Zhang +4

Warning: This paper contains examples of harmful language, and reader discretion is recommended. The increasing open release of powerful large language models (LLMs) has facilitate…