activity
20242026
collaborators

5 papers

cs.AI2026

A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models

Xiangyu Yin, Tora Bodin, Rohan Menon +1

Inference time defences against vision language model jailbreaks often subtract a calibrated direction from the residual stream at a chosen decoder layer. We compare five defence c…

cs.AI2026

Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration

Xiangyu Yin, Jiaxu Liu, Zhen Chen +1

Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tuning on twenty forget examples substantially recovers held-out forget-set ROUGE for e…

cs.AI2026

ProGRank: Probe-Gradient Reranking to Defend Dense-Retriever RAG from Corpus Poisoning

Xiangyu Yin, Yi Qi, Chih-Hong Cheng

Retrieval-Augmented Generation (RAG) improves large language model applications by grounding generation in retrieved evidence, but also introduces corpus poisoning as a new attack…

cs.LG2025

Randomized Smoothing Meets Vision-Language Models

Emmanouil Seferis, Changshun Wu, Stefanos Kollias +2

Randomized smoothing (RS) is one of the prominent techniques to ensure the correctness of machine learning models, where point-wise robustness certificates can be derived analytica…

cs.LG2024

Estimating the Robustness Radius for Randomized Smoothing with 100 Sample Efficiency

Emmanouil Seferis, Stefanos Kollias, Chih-Hong Cheng

Randomized smoothing (RS) has successfully been used to improve the robustness of predictions for deep neural networks (DNNs) by adding random noise to create multiple variations o…