collaborators

40 papers

cs.CL2026

Safety Cost of Steering Vectors Is Separable and Reducible

Yuxiao Li, Gjergji Kasneci

Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally compromise a model's safety mechani…

cs.LG2026

Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning

Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj +1

The paper investigates how different reward functions affect the speed and effectiveness of reinforcement‑learning based machine unlearning for language models, proposing graded an…

cs.CY2026

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

Gjergji Kasneci, Enkelejda Kasneci

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is i…

cs.AI2026

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45

AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…

cs.RO2026

Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients

Carmen Ng, Gjergji Kasneci

LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms vary across cultures by age, status, and group size, failure to cal…

cs.CL2026

Consolidating Rewarded Perturbations for LLM Post-Training

Zheyu Zhang, Shuo Yang, Gjergji Kasneci

Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by RandOpt, relocates this loo…