12 papers
Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness
Haoting Qian, Qingjie Zhang, Zhicong Huang +2
Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current unlearning benchmarks inc…
DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models
Zhicong Huang, Cheng Hong, Tao Wei
Cloud-based language model services routinely process prompts containing sensitive information. Obfuscation-based defenses---including ObfusLM, SentinelLMs, TextObfuscator, and DPN…
When New Generators Arrive: Lifelong Machine-Generated Text Attribution via Ridge Feature Transfer
Zhen Sun, Yifan Liao, Zhicong Huang +4
Machine-generated text (MGT) attribution aims to identify the specific generator responsible for a given text, thereby providing fine-grained evidence for model accountability and…
Fingerprinting LLMs via Prompt Injection
Yuepeng Hu, Zhengyuan Jiang, Mengyuan Li +4
Large language models (LLMs) are often modified after release through post-processing such as post-training or quantization, which makes it challenging to determine whether one mod…
FIT to Forget: Robust Continual Unlearning for Large Language Models
Xiaoyu Xu, Minxin Du, Kun Fang +5
While large language models (LLMs) exhibit remarkable capabilities, they increasingly face demands to unlearn memorized privacy-sensitive, copyrighted, or harmful content. Existing…
VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts
Peigui Qi, Kunsheng Tang, Yanpu Yu +7
Vision-Language Models (VLMs) face significant safety vulnerabilities from malicious prompt attacks due to weakened alignment during visual integration. Existing defenses suffer fr…