3 papers
cs.LG2025
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
Filip Sondej, Yushi Yang, MikoÅaj Kniejski +1
Language models can retain dangerous knowledge and skills even after extensive safety fine-tuning, posing both misuse and misalignment risks. Recent studies show that even speciali…
cs.CV2025
Personalized Interpretability -- Interactive Alignment of Prototypical Parts Networks
Tomasz Michalski, Adam Wróbel, Andrea Bontempelli +6
Concept-based interpretable neural networks have gained significant attention due to their intuitive and easy-to-understand explanations based on case-based reasoning, such as "thi…
cs.AI2025
Multi-Agent Security Tax: Trading Off Security and Collaboration Capabilities in Multi-Agent Systems
Pierre Peigne-Lefebvre, Mikolaj Kniejski, Filip Sondej +4
As AI agents are increasingly adopted to collaborate on complex objectives, ensuring the security of autonomous multi-agent systems becomes crucial. We develop simulations of agent…