2 papers
cs.AI2026
Statistical Estimation of Adversarial Risk in Large Language Models under Best-of-N Sampling
Mingqian Feng, Xiaodong Liu, Weiwei Yang +3
Large Language Models (LLMs) are typically evaluated for safety under single-shot or low-budget adversarial prompting, which underestimates real-world risk. In practice, attackers…
cs.CL2026
SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks
Mingqian Feng, Xiaodong Liu, Weiwei Yang +4
Multi-turn jailbreaks capture the real threat model for safety-aligned chatbots, where single-turn attacks are merely a special case. Yet existing approaches break under exploratio…