5 papers
A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models
Xiangyu Yin, Tora Bodin, Rohan Menon +1
Inference time defences against vision language model jailbreaks often subtract a calibrated direction from the residual stream at a chosen decoder layer. We compare five defence c…
Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration
Xiangyu Yin, Jiaxu Liu, Zhen Chen +1
Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tuning on twenty forget examples substantially recovers held-out forget-set ROUGE for e…
ProGRank: Probe-Gradient Reranking to Defend Dense-Retriever RAG from Corpus Poisoning
Xiangyu Yin, Yi Qi, Chih-Hong Cheng
Retrieval-Augmented Generation (RAG) improves large language model applications by grounding generation in retrieved evidence, but also introduces corpus poisoning as a new attack…
Randomized Smoothing Meets Vision-Language Models
Emmanouil Seferis, Changshun Wu, Stefanos Kollias +2
Randomized smoothing (RS) is one of the prominent techniques to ensure the correctness of machine learning models, where point-wise robustness certificates can be derived analytica…
Estimating the Robustness Radius for Randomized Smoothing with 100 Sample Efficiency
Emmanouil Seferis, Stefanos Kollias, Chih-Hong Cheng
Randomized smoothing (RS) has successfully been used to improve the robustness of predictions for deep neural networks (DNNs) by adding random noise to create multiple variations o…