2 papers
cs.CL2025
On the Loss of Context-awareness in General Instruction Fine-tuning
Yihan Wang, Andrew Bai, Nanyun Peng +1
Pre-trained Large Language Models (LLMs) require post-training methods such as supervised fine-tuning (SFT) on instruction-response pairs to enable instruction following. However,…
cs.CL2024
Defending LLMs against Jailbreaking Attacks via Backtranslation
Yihan Wang, Zhouxing Shi, Andrew Bai +1
Although many large language models (LLMs) have been trained to refuse harmful requests, they are still vulnerable to jailbreaking attacks which rewrite the original prompt to conc…