Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Low-Cost Hard-Label Adversarial Attack with Theoretical Foundations
Jun Liu, Leo Yu Zhang, Fengpeng Li +2
Hard-label black-box attacks, relying solely on top-1 predictions, represent one of the most challenging yet practically threat models. Despite recent progress, existing approaches…
cs.LG2026
LLM Unlearning with LLM Beliefs
Kemou Li, Qizhou Wang, Yue Wang +4
Large language models trained on vast corpora inherently risk memorizing sensitive or harmful content, which may later resurface in their outputs. Prevailing unlearning methods gen…