2 papers
cs.CL2025
Unnatural Languages Are Not Bugs but Features for LLMs
Keyu Duan, Yiran Zhao, Zhili Feng +9
Large Language Models (LLMs) have been observed to process non-human-readable text sequences, such as jailbreak prompts, often viewed as a bug for aligned LLMs. In this work, we pr…
cs.CL2024
Accelerating Greedy Coordinate Gradient and General Prompt Optimization via Probe Sampling
Yiran Zhao, Wenyue Zheng, Tianle Cai +4
Safety of Large Language Models (LLMs) has become a critical issue given their rapid progresses. Greedy Coordinate Gradient (GCG) is shown to be effective in constructing adversari…