Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models
Jiahui Li, Yongchang Hao, Haoyu Xu +2
Despite the advancements in training Large Language Models (LLMs) with alignment techniques to enhance the safety of generated content, these models remain susceptible to jailbreak…
cs.CL2024
Evaluating Knowledge-based Cross-lingual Inconsistency in Large Language Models
Xiaolin Xing, Zhiwei He, Haoyu Xu +3
This paper investigates the cross-lingual inconsistencies observed in Large Language Models (LLMs), such as ChatGPT, Llama, and Baichuan, which have shown exceptional performance i…