4 papers
Analysing the Safety Pitfalls of Steering Vectors
Yuxiao Li, Alina Fastowski, Efstratios Zaradoukas +2
Activation steering has emerged as a powerful tool to shape LLM behavior without the need for weight updates. While its inherent brittleness and unreliability are well-documented,…
Reinforcement Unlearning via Group Relative Policy Optimization
Efstratios Zaradoukas, Bardh Prenkaj, Gjergji Kasneci
During pretraining, LLMs inadvertently memorize sensitive or copyrighted data, posing significant compliance challenges under legal frameworks like the GDPR and the EU AI Act. Fulf…
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
Mahdi Dhaini, Kafaite Zahra Hussain, Efstratios Zaradoukas +1
As Natural Language Processing (NLP) models continue to evolve and become integral to high-stakes applications, ensuring their interpretability remains a critical challenge. Given…
Graph Inverse Style Transfer for Counterfactual Explainability
Bardh Prenkaj, Efstratios Zaradoukas, Gjergji Kasneci
Counterfactual explainability seeks to uncover model decisions by identifying minimal changes to the input that alter the predicted outcome. This task becomes particularly challeng…