most citedShadow Alignment: The Ease of Subverting Safely-Aligned Language Models

10 citations · 15 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2023

Democratizing Reasoning Ability: Tailored Learning from Large Language Model

Zhaoyang Wang, Shaohan Huang, Yuxuan Liu +8

Large language models (LLMs) exhibit impressive emergent abilities in natural language processing, but their democratization is hindered due to huge computation requirements and cl…

cs.CL20234 cited

TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models

Xiao Wang, Yuansen Zhang, Tianze Chen +9

Aligned large language models (LLMs) demonstrate exceptional capabilities in task-solving, following instructions, and ensuring safety. However, the continual learning aspect of th…

cs.CV20231 cited

Towards Domain-Specific Features Disentanglement for Domain Generalization

Hao Chen, Qi Zhang, Zenan Huang +2

Distributional shift between domains poses great challenges to modern machine learning algorithms. The domain generalization (DG) signifies a popular line targeting this issue, whe…

cs.CL202310 cited

Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Xianjun Yang, Xiao Wang, Qi Zhang +4

Warning: This paper contains examples of harmful language, and reader discretion is recommended. The increasing open release of powerful large language models (LLMs) has facilitate…

cs.IR2023

Model-enhanced Vector Index

Hailin Zhang, Yujing Wang, Qi Chen +16

Embedding-based retrieval methods construct vector indices to search for document representations that are most similar to the query representations. They are widely used in docume…