activity
20232026
most citedShadow Alignment: The Ease of Subverting Safely-Aligned Language Models

10 citations · 12 across the 5 of their papers we have counts for

collaborators

6 papers

q-bio.BM2026

AMix-2: Establishing Protein as a Native Modality in Large Language Models

Keyue Qiu, Yixin Wu, Lihao Wang +19

We present AMix-2, a protein-text foundation model that establishes protein as a native modality in large language models (LLMs), unifying protein understanding and sequence design…

cs.CL20241 cited

Uncertainty Aware Learning for Language Model Alignment

Yikun Wang, Rui Zheng, Liang Ding +3

As instruction-tuned large language models (LLMs) evolve, aligning pretrained foundation models presents increasing challenges. Existing alignment strategies, which typically lever…

cs.CL20241 cited

Unveiling the Misuse Potential of Base Large Language Models via In-Context Learning

Xiao Wang, Tianze Chen, Xianjun Yang +3

The open-sourcing of large language models (LLMs) accelerates application development, innovation, and scientific progress. This includes both base models, which are pre-trained on…

cs.CL2024

Navigating the OverKill in Large Language Models

Chenyu Shi, Xiao Wang, Qiming Ge +7

Large language models are meticulously aligned to be both helpful and harmless. However, recent research points to a potential overkill which means models may refuse to answer beni…

cs.CL2024

Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback

Songyang Gao, Qiming Ge, Wei Shen +9

The success of AI assistants based on Language Models (LLMs) hinges on Reinforcement Learning from Human Feedback (RLHF) to comprehend and align with user intentions. However, trad…

cs.CL202310 cited

Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Xianjun Yang, Xiao Wang, Qi Zhang +4

Warning: This paper contains examples of harmful language, and reader discretion is recommended. The increasing open release of powerful large language models (LLMs) has facilitate…