9 papers
ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models
Hashmat Shadab Malik, Toluwani Aremu, Samuele Poppi +2
Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However,…
SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation
Mohammed Talha Alam, Nada Saadi, Fahad Shamshad +4
Text-to-image diffusion models can emit copyrighted, unsafe, or private content. Safety alignment aims to suppress specific concepts, yet evaluations seldom test whether safety per…
A Gravitational Interpretation of Fine-Tuning Reversion
Samuele Poppi, Nils Lukas
Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearned capabilities can re-emerge,…
When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
Sai Kartheek Reddy Kasu, Nils Lukas, Samuele Poppi
Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in a long dialogue, yet its final-turn refu…
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
Toluwani Aremu, Noor Hussein, Munachiso Nwadike +5
Watermarking enables GenAI providers to verify whether content was generated by their models. A watermark is a hidden signal in the content, whose presence can be detected using a…
Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
Silvia Cappelletti, Tobia Poppi, Samuele Poppi +5
Large Language Models (LLMs) are increasingly evaluated on multiple-choice question answering (MCQA) tasks using *first-token probability* (FTP), which selects the answer option wh…