4 papers
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
Benjamin Arnav, Pablo Bernabeu-Pérez, Nathan Helm-Burger +3
As AI models are deployed with increasing autonomy, it is important to ensure they do not take harmful actions unnoticed. As a potential mitigation, we investigate Chain-of-Thought…
The Aloe Family Recipe for Open and Specialized Healthcare LLMs
Dario Garcia-Gasulla, Jordi Bayarri-Planas, Ashwin Kumar Gururajan +10
Purpose: With advancements in Large Language Models (LLMs) for healthcare, the need arises for competitive open-source models to protect the public interest. This work contributes…
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
Anna Arias-Duart, Pablo Agustin Martin-Torres, Daniel Hinjos +7
Current Large Language Models (LLMs) benchmarks are often based on open-ended or close-ended QA evaluations, avoiding the requirement of human labor. Close-ended measurements evalu…
Present and Future Generalization of Synthetic Image Detectors
Pablo Bernabeu-Perez, Enrique Lopez-Cuena, Dario Garcia-Gasulla
The continued release of increasingly realistic image generation models creates a demand for synthetic image detectors. To build effective detectors we must first understand how fa…