collaborators

12 papers

cs.AI2026

Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness

Haoting Qian, Qingjie Zhang, Zhicong Huang +2

Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current unlearning benchmarks inc…

cs.CR2026

DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models

Zhicong Huang, Cheng Hong, Tao Wei

Cloud-based language model services routinely process prompts containing sensitive information. Obfuscation-based defenses---including ObfusLM, SentinelLMs, TextObfuscator, and DPN…

cs.CL2026

When New Generators Arrive: Lifelong Machine-Generated Text Attribution via Ridge Feature Transfer

Zhen Sun, Yifan Liao, Zhicong Huang +4

Machine-generated text (MGT) attribution aims to identify the specific generator responsible for a given text, thereby providing fine-grained evidence for model accountability and…

cs.CR2026

Fingerprinting LLMs via Prompt Injection

Yuepeng Hu, Zhengyuan Jiang, Mengyuan Li +4

Large language models (LLMs) are often modified after release through post-processing such as post-training or quantization, which makes it challenging to determine whether one mod…

cs.CL2026

FIT to Forget: Robust Continual Unlearning for Large Language Models

Xiaoyu Xu, Minxin Du, Kun Fang +5

While large language models (LLMs) exhibit remarkable capabilities, they increasingly face demands to unlearn memorized privacy-sensitive, copyrighted, or harmful content. Existing…

cs.LG2026

VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts

Peigui Qi, Kunsheng Tang, Yanpu Yu +7

Vision-Language Models (VLMs) face significant safety vulnerabilities from malicious prompt attacks due to weakened alignment during visual integration. Existing defenses suffer fr…