5 papers
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions
Yukai Zhou, Feiyang Lu, Xiaokai Mao +2
Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily…
DaDaDa: A Dataset for Data Pricing in Data Marketplaces
Qiheng Sun, Hongwei Zhang, Junxu Liu +4
High-quality data drives machine learning advances across industries. Recognizing the value of data, data transactions are increasingly common, giving rise to many data marketplace…
Feature Attribution in Directed Acyclic Graphs Using Edge Intervention
Qiheng Sun, Junxu Liu, Xiaokai Mao +4
Shapley value-based feature attribution methods face challenges in scenarios involving complex feature interactions and causal relationships, even when a causal structure is provid…
CoKV: Optimizing KV Cache Allocation via Cooperative Game
Qiheng Sun, Hongwei Zhang, Haocheng Xia +3
Large language models (LLMs) have achieved remarkable success on various aspects of human life. However, one of the major challenges in deploying these models is the substantial me…
Prompt Valuation Based on Shapley Values
Hanxi Liu, Xiaokai Mao, Haocheng Xia +3
Large language models (LLMs) excel on new tasks without additional training, simply by providing natural language prompts that demonstrate how the task should be performed. Prompt…