Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
Privacy Risks and Preservation Methods in Explainable Artificial Intelligence: A Scoping Review
Sonal Allana, Mohan Kankanhalli, Rozita Dara
Explainable Artificial Intelligence (XAI) has emerged as a pillar of Trustworthy AI and aims to bring transparency in complex models that are opaque by nature. Despite the benefits…
cs.AI2025
Bullying the Machine: How Personas Increase LLM Vulnerability
Ziwei Xu, Udit Sanghi, Mohan Kankanhalli
Large Language Models (LLMs) are increasingly deployed in interactions where they are prompted to adopt personas. This paper investigates whether such persona conditioning affects…
cs.AI2025
Strong Preferences Affect the Robustness of Preference Models and Value Alignment
Ziwei Xu, Mohan Kankanhalli
Value alignment, which aims to ensure that large language models (LLMs) and other AI agents behave in accordance with human values, is critical for ensuring safety and trustworthin…