28 papers
Zoom In Disparities in Healthcare LLM Q&A
Ipek Baris Schlicht, Burcu Sayin, Zhixue Zhao +5
Equitable access to reliable health information is vital when integrating AI into healthcare. Yet, information quality varies across languages, raising concerns about the reliabili…
Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows
Shivam Rawat, Lucie Flek
Agentic AI systems are increasingly being integrated into scientific workflows, yet their behavior under realistic conditions remains insufficiently understood. We evaluate CMBAgen…
MiniFool -- Physics-Constraint-Aware Minimizer-Based Adversarial Attacks in Deep Neural Networks
Lucie Flek, Oliver Janik, Philipp Alexander Jung +8
In this paper, we present a new algorithm, MiniFool, that implements physics-inspired adversarial attacks for testing neural network-based classification tasks in particle and astr…
Personality Anchoring for Social Simulation: Linking Personality, Social Behavior, and Interaction Success with LLM Agents
Vahid Sadiri Javadi, Aksa Aksa, Fryderyk Róg +2
Social interactions are shaped by the interplay of dispositional traits and situational context, yet systematically investigating how personality configurations between individuals…
Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
Magnus Jørgenvåg, David Kaczér, Lasse Ruttert +3
Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples. While EM has been extensive…
Shapes are not enough: CONSERVAttack and its use for finding vulnerabilities and uncertainties in machine learning applications
Philip Bechtle, Lucie Flek, Philipp Alexander Jung +7
In High Energy Physics, as in many other fields of science, the application of machine learning techniques has been crucial in advancing our understanding of fundamental phenomena.…