Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
LLM Benchmark Datasets Should Be Contamination-Resistant
Ali Al-Lawati, Jason Lucas, Dongwon Lee +1
Benchmark datasets are critical for reproducible, reliable, and discriminative evaluation of LLMs. However, recent studies reveal that many benchmark datasets are included in pretr…
cs.LG2025
SUA: Stealthy Multimodal Large Language Model Unlearning Attack
Xianren Zhang, Hui Liu, Delvin Ce Zhang +4
Multimodal Large Language Models (MLLMs) trained on massive data may memorize sensitive personal information and photos, posing serious privacy risks. To mitigate this, MLLM unlear…