large language models 1multi-turn jailbreak 1policy optimization 1reinforcement learning 1turn-level credit assignment 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment
Junyoung Park, Namgyu Park, Sechan Lee +3
The paper proposes a turn-level credit assignment framework (DC‑GRPO) for training multi‑turn jailbreak attacks on large language models, showing higher success rates than prior me…
cs.CR2026
Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
Sechan Lee, Hyounghun Kim, Sangdon Park
Malicious AI causing harm to humans is not just a Hollywood fantasy. Indeed, as highly capable models such as Claude Mythos emerge and agent systems like OpenClaw rapidly spread, t…