From the 1 of 5 linked papers with an AI index.
5 papers
MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment
Junyoung Park, Namgyu Park, Sechan Lee +3
The paper proposes a turn-level credit assignment framework (DC‑GRPO) for training multi‑turn jailbreak attacks on large language models, showing higher success rates than prior me…
The Interplay of Harness Design and Post-Training in LLM Agents
Kyungmin Kim, Youngbin Choi, Seoyeon Lee +3
Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and what auxiliary information accom…
Online Conformal Abstention for Factuality Control Under Adversarial Bandit Feedback
Minjae Lee, Yoonjae Jung, Sangdon Park
As interactive generative systems are increasingly deployed in real-world applications, their tendency to generate unreliable or false responses raises serious concerns. Conformal…
Transductive Generalization via Optimal Transport and Its Application to Graph Node Classification
MoonJeong Park, Seungbeom Lee, Kyungmin Kim +5
Many existing transductive bounds rely on classical complexity measures that are computationally intractable and often misaligned with empirical behavior. In this work, we establis…
Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model Unlearning
Saemi Moon, Minjong Lee, Sangdon Park +1
As text-to-image diffusion models gain widespread commercial applications, there are increasing concerns about unethical or harmful use, including the unauthorized generation of co…