4 papers · 1 filter
Shieldstral
Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +273
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7 its size on text safety benchmarks and set…
On the Loss of Context-awareness in General Instruction Fine-tuning
Yihan Wang, Andrew Bai, Nanyun Peng +1
Pre-trained Large Language Models (LLMs) require post-training methods such as supervised fine-tuning (SFT) on instruction-response pairs to enable instruction following. However,…
CLUE: Concept-Level Uncertainty Estimation for Large Language Models
Yu-Hsiang Wang, Andrew Bai, Che-Ping Tsai +1
Large Language Models (LLMs) have demonstrated remarkable proficiency in various natural language generation (NLG) tasks. Previous studies suggest that LLMs' generation process inv…
Defending LLMs against Jailbreaking Attacks via Backtranslation
Yihan Wang, Zhouxing Shi, Andrew Bai +1
Although many large language models (LLMs) have been trained to refuse harmful requests, they are still vulnerable to jailbreaking attacks which rewrite the original prompt to conc…