13 papers
Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness
Haoting Qian, Qingjie Zhang, Zhicong Huang +2
Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current unlearning benchmarks inc…
When New Generators Arrive: Lifelong Machine-Generated Text Attribution via Ridge Feature Transfer
Zhen Sun, Yifan Liao, Zhicong Huang +4
Machine-generated text (MGT) attribution aims to identify the specific generator responsible for a given text, thereby providing fine-grained evidence for model accountability and…
Turn-Based Structural Triggers: Structure-Conditioned Backdoors in Multi-Turn LLMs
Yiyang Lu, Jinwen He, Yue Zhao +4
Large Language Models (LLMs) are increasingly deployed as multi-turn assistants and customized through instruction tuning with project-specific training components. This practice c…
Fingerprinting LLMs via Prompt Injection
Yuepeng Hu, Zhengyuan Jiang, Mengyuan Li +4
Large language models (LLMs) are often modified after release through post-processing such as post-training or quantization, which makes it challenging to determine whether one mod…
FIT to Forget: Robust Continual Unlearning for Large Language Models
Xiaoyu Xu, Minxin Du, Kun Fang +5
While large language models (LLMs) exhibit remarkable capabilities, they increasingly face demands to unlearn memorized privacy-sensitive, copyrighted, or harmful content. Existing…
Robustness of Vision Foundation Models to Common Perturbations
Hongbin Liu, Zhengyuan Jiang, Cheng Hong +1
A vision foundation model outputs an embedding vector for an image, which can be affected by common editing operations (e.g., JPEG compression, brightness, contrast adjustments). T…