From the 1 of 5 linked papers with an AI index.
5 papers
Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning
Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj +1
The paper investigates how different reward functions affect the speed and effectiveness of reinforcement‑learning based machine unlearning for language models, proposing graded an…
Analysing the Safety Pitfalls of Steering Vectors
Yuxiao Li, Alina Fastowski, Efstratios Zaradoukas +2
Activation steering has emerged as a powerful tool to shape LLM behavior without the need for weight updates. While its inherent brittleness and unreliability are well-documented,…
Reinforcement Unlearning via Group Relative Policy Optimization
Efstratios Zaradoukas, Bardh Prenkaj, Gjergji Kasneci
During pretraining, LLMs inadvertently memorize sensitive or copyrighted data, posing significant compliance challenges under legal frameworks like the GDPR and the EU AI Act. Fulf…
Graph Inverse Style Transfer for Counterfactual Explainability
Bardh Prenkaj, Efstratios Zaradoukas, Gjergji Kasneci
Counterfactual explainability seeks to uncover model decisions by identifying minimal changes to the input that alter the predicted outcome. This task becomes particularly challeng…
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
Mahdi Dhaini, Kafaite Zahra Hussain, Efstratios Zaradoukas +1
As Natural Language Processing (NLP) models continue to evolve and become integral to high-stakes applications, ensuring their interpretability remains a critical challenge. Given…