activity
20242026
collaborators

7 papers

cs.CR2026

Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense

Minseok Choi, Seungbin Yang, Dongjin Kim +5

Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifiers cannot adapt to these evol…

cs.CL2026

ExpGuard: LLM Content Moderation in Specialized Domains

Minseok Choi, Dongjin Kim, Seungbin Yang +5

With the growing deployment of large language models (LLMs) in real-world applications, establishing robust safety guardrails to moderate their inputs and outputs has become essent…

cs.CL2025

Opt-Out: Investigating Entity-Level Unlearning for Large Language Models via Optimal Transport

Minseok Choi, Daniel Rim, Dohyun Lee +1

Instruction-following large language models (LLMs), such as ChatGPT, have become widely popular among everyday users. However, these models inadvertently disclose private, sensitiv…

cs.CL2024

Breaking Chains: Unraveling the Links in Multi-Hop Knowledge Unlearning

Minseok Choi, ChaeHun Park, Dohyun Lee +1

Large language models (LLMs) serve as giant information stores, often including personal or copyrighted data, and retraining them from scratch is not a viable option. This has led…

cs.CL2024

Cross-Lingual Unlearning of Selective Knowledge in Multilingual Language Models

Minseok Choi, Kyunghyun Min, Jaegul Choo

Pretrained language models memorize vast amounts of information, including private and copyrighted data, raising significant safety concerns. Retraining these models after excludin…

cs.CL2024

PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison

ChaeHun Park, Minseok Choi, Dohyun Lee +1

Building a reliable and automated evaluation metric is a necessary but challenging problem for open-domain dialogue systems. Recent studies proposed evaluation metrics that assess…