From the 1 of 4 linked papers with an AI index.
4 papers
HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing
Euntae Kim, Soomin Han, Buru Chang
The paper introduces HarDBench, a benchmark that evaluates how vulnerable large language models are to jailbreak attacks when used as co-authors in draft-based writing, and propose…
Alignment Data Map for Efficient Preference Data Selection and Diagnosis
Seohyeong Lee, Eunwon Kim, Hwaran Lee +1
Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and inefficient-motivating the need for eff…
Improving Fisher Information Estimation and Efficiency for LoRA-based LLM Unlearning
Yejin Kim, Eunwon Kim, Buru Chang +1
LLMs have demonstrated remarkable performance across various tasks but face challenges related to unintentionally generating outputs containing sensitive information. A straightfor…
SHARE: Shared Memory-Aware Open-Domain Long-Term Dialogue Dataset Constructed from Movie Script
Eunwon Kim, Chanho Park, Buru Chang
Shared memories between two individuals strengthen their bond and are crucial for facilitating their ongoing conversations. This study aims to make long-term dialogue more engaging…