4 papers
Language Models Can Explain Visual Features via Steering
Javier Ferrando, Enrique Lopez-Cuena, Pablo Agustin Martin-Torres +3
Sparse Autoencoders uncover thousands of features in vision models, yet explaining these features without requiring human intervention remains an open challenge. While previous wor…
The Aloe Family Recipe for Open and Specialized Healthcare LLMs
Dario Garcia-Gasulla, Jordi Bayarri-Planas, Ashwin Kumar Gururajan +10
Purpose: With advancements in Large Language Models (LLMs) for healthcare, the need arises for competitive open-source models to protect the public interest. This work contributes…
Efficient Safety Retrofitting Against Jailbreaking for LLMs
Dario Garcia-Gasulla, Adrian Tormos, Anna Arias-Duart +4
Direct Preference Optimization (DPO) is an efficient alignment technique that steers LLMs towards preferable outputs by training on preference data, bypassing the need for explicit…
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
Anna Arias-Duart, Pablo Agustin Martin-Torres, Daniel Hinjos +7
Current Large Language Models (LLMs) benchmarks are often based on open-ended or close-ended QA evaluations, avoiding the requirement of human labor. Close-ended measurements evalu…