Showing cs.CYShow all
2 papers · 1 filter
cs.CY2026
Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
Naveen Raman, Santiago Cortes-Gomez, Mateo Dulce Rubio +2
Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not bec…
cs.CY2025
SESGO: Spanish Evaluation of Stereotypical Generative Outputs
Melissa Robles, Catalina Bernal, Denniss Raigoso +1
This paper addresses the critical gap in evaluating bias in multilingual Large Language Models (LLMs), with a specific focus on Spanish language within culturally-aware Latin Ameri…