27 citations · 142 across the 45 of their papers we have counts for
68 papers
When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness
Mahdi Dhaini, Adam Dejl, Juraj Vladika +3
Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in in-context learning (ICL). However, it re…
Safety Cost of Steering Vectors Is Separable and Reducible
Yuxiao Li, Gjergji Kasneci
Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally compromise a model's safety mechani…
Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning
Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj +1
Machine unlearning seeks to selectively remove specific knowledge from trained language models without full retraining, a growing necessity under privacy regulations such as GDPR a…
The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
Gjergji Kasneci, Enkelejda Kasneci
Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is i…
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients
Carmen Ng, Gjergji Kasneci
LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms vary across cultures by age, status, and group size, failure to cal…