6 papers
PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement
Xianren Zhang, Delvin Ce Zhang, Dongwon Lee +1
Multimodal Large Language Models (MLLMs) have shown strong capabilities, but they may memorize private information from web data, raising privacy concerns. Machine unlearning offer…
LLM Benchmark Datasets Should Be Contamination-Resistant
Ali Al-Lawati, Jason Lucas, Dongwon Lee +1
Benchmark datasets are critical for reproducible, reliable, and discriminative evaluation of LLMs. However, recent studies reveal that many benchmark datasets are included in pretr…
Query-Efficient Agentic Graph Extraction Attacks on GraphRAG Systems
Shuhua Yang, Jiahao Zhang, Yilong Wang +2
Graph-based retrieval-augmented generation (GraphRAG) systems construct knowledge graphs over document collections to support multi-hop reasoning. While prior work shows that Graph…
SUA: Stealthy Multimodal Large Language Model Unlearning Attack
Xianren Zhang, Hui Liu, Delvin Ce Zhang +4
Multimodal Large Language Models (MLLMs) trained on massive data may memorize sensitive personal information and photos, posing serious privacy risks. To mitigate this, MLLM unlear…
Divide-Verify-Refine: Can LLMs Self-Align with Complex Instructions?
Xianren Zhang, Xianfeng Tang, Hui Liu +4
Recent studies show LLMs struggle with complex instructions involving multiple constraints (e.g., length, format, sentiment). Existing works address this issue by fine-tuning, whic…
CORRECT: Context- and Reference-Augmented Reasoning and Prompting for Fact-Checking
Delvin Ce Zhang, Dongwon Lee
Fact-checking the truthfulness of claims usually requires reasoning over multiple evidence sentences. Oftentimes, evidence sentences may not be always self-contained, and may requi…