40 papers
Safety Cost of Steering Vectors Is Separable and Reducible
Yuxiao Li, Gjergji Kasneci
Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally compromise a model's safety mechani…
Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning
Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj +1
The paper investigates how different reward functions affect the speed and effectiveness of reinforcement‑learning based machine unlearning for language models, proposing graded an…
The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
Gjergji Kasneci, Enkelejda Kasneci
Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is i…
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients
Carmen Ng, Gjergji Kasneci
LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms vary across cultures by age, status, and group size, failure to cal…
Consolidating Rewarded Perturbations for LLM Post-Training
Zheyu Zhang, Shuo Yang, Gjergji Kasneci
Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by RandOpt, relocates this loo…