4 papers
Towards the Effect of Examples on In-Context Learning: A Theoretical Case Study
Pengfei He, Yingqian Cui, Han Xu +4
In-context learning (ICL) has emerged as a powerful capability for large language models (LLMs) to adapt to downstream tasks by leveraging a few (demonstration) examples. Despite i…
Data Poisoning for In-context Learning
Pengfei He, Han Xu, Yue Xing +3
In the domain of large language models (LLMs), in-context learning (ICL) has been recognized for its innovative ability to adapt to new tasks, relying on examples rather than retra…
Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis
Yuping Lin, Pengfei He, Han Xu +4
Large language models (LLMs) are susceptible to a type of attack known as jailbreaking, which misleads LLMs to output harmful contents. Although there are diverse jailbreak attack…
Stealthy Backdoor Attack via Confidence-driven Sampling
Pengfei He, Yue Xing, Han Xu +6
Backdoor attacks aim to surreptitiously insert malicious triggers into DNN models, granting unauthorized control during testing scenarios. Existing methods lack robustness against…