works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.LG2026

Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning

Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj +1

The paper investigates how different reward functions affect the speed and effectiveness of reinforcement‑learning based machine unlearning for language models, proposing graded an…

cs.CR2026

Analysing the Safety Pitfalls of Steering Vectors

Yuxiao Li, Alina Fastowski, Efstratios Zaradoukas +2

Activation steering has emerged as a powerful tool to shape LLM behavior without the need for weight updates. While its inherent brittleness and unreliability are well-documented,…

cs.LG2026

Reinforcement Unlearning via Group Relative Policy Optimization

Efstratios Zaradoukas, Bardh Prenkaj, Gjergji Kasneci

During pretraining, LLMs inadvertently memorize sensitive or copyrighted data, posing significant compliance challenges under legal frameworks like the GDPR and the EU AI Act. Fulf…

cs.LG2025

Graph Inverse Style Transfer for Counterfactual Explainability

Bardh Prenkaj, Efstratios Zaradoukas, Gjergji Kasneci

Counterfactual explainability seeks to uncover model decisions by identifying minimal changes to the input that alter the predicted outcome. This task becomes particularly challeng…

cs.CL2025

EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models

Mahdi Dhaini, Kafaite Zahra Hussain, Efstratios Zaradoukas +1

As Natural Language Processing (NLP) models continue to evolve and become integral to high-stakes applications, ensuring their interpretability remains a critical challenge. Given…