2 papers
cs.AI2025
Bias by Design? How Data Practices Shape Fairness in AI Healthcare Systems
Anna Arias-Duart, Maria Eugenia Cardello, Atia Cortés
Artificial intelligence (AI) holds great promise for transforming healthcare. However, despite significant advances, the integration of AI solutions into real-world clinical practi…
cs.CL2025
Efficient Safety Retrofitting Against Jailbreaking for LLMs
Dario Garcia-Gasulla, Adrian Tormos, Anna Arias-Duart +4
Direct Preference Optimization (DPO) is an efficient alignment technique that steers LLMs towards preferable outputs by training on preference data, bypassing the need for explicit…