16 citations · 16 across the 2 of their papers we have counts for
2 papers
cs.CY2026
Auditing Alignment Controllability in LLMs via Political Axes
Bartol Bućan, Nikola Sočec, Sarah Isufi +5
Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land som…
cs.AI2022★ 16 cited
Explainability in reinforcement learning: perspective and position
Agneza Krajna, Mario Brcic, Tomislav Lipic +1
Artificial intelligence (AI) has been embedded into many aspects of people's daily lives and it has become normal for people to have AI make decisions for them. Reinforcement learn…