7 citations · 9 across the 8 of their papers we have counts for
8 papers
Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions
Minh Duc Bui, Mario Sanz-Guerrero, Abteen Ebrahimi +4
Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to cultural cues? Existing multilingual medical benchmarks usually assume…
Large Language Models Are Overconfident in Their Own Responses
Mario Sanz-Guerrero, Manuel Mager, Katharina von der Wense
Prior work has shown that instruction-tuned large language models (LLMs) are less well calibrated than their base pre-trained counterparts. However, little is known about the frequ…
NALA_MAINZ at BLP-2025 Task 2: A Multi-agent Approach for Bangla Instruction to Python Code Generation
Hossain Shaikh Saadi, Faria Alam, Mario Sanz-Guerrero +3
This paper presents JGU Mainz's winning system for the BLP-2025 Shared Task on Code Generation from Bangla Instructions. We propose a multi-agent-based pipeline. First, a code-gene…
Mitigating Label Length Bias in Large Language Models
Mario Sanz-Guerrero, Katharina von der Wense
Large language models (LLMs) are powerful zero- and few-shot learners. However, when predicting over a set of candidate options, LLMs suffer from label biases, and existing calibra…
JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA
Hossain Shaikh Saadi, Minh Duc Bui, Mario Sanz-Guerrero +1
This paper presents the JGU Mainz submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: Machine Translation and Question Answering, focusing on U…
Corrective In-Context Learning: Evaluating Self-Correction in Large Language Models
Mario Sanz-Guerrero, Katharina von der Wense
In-context learning (ICL) has transformed the use of large language models (LLMs) for NLP tasks, enabling few-shot learning by conditioning on labeled examples without finetuning.…