2 papers
cs.CL2026
EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
Feitong Qiao, Liren Peng, Shiming Ren +7
Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least…
cs.AI2026
Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation
Genglin Liu, Muye Zhang, Krishnamurthy Viswanathan +5
Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-dis…