4 papers
Large Language Models Are Overconfident in Their Own Responses
Mario Sanz-Guerrero, Manuel Mager, Katharina von der Wense
Prior work has shown that instruction-tuned large language models (LLMs) are less well calibrated than their base pre-trained counterparts. However, little is known about the frequ…
From If-Statements to ML Pipelines: Revisiting Bias in Code-Generation
Minh Duc Bui, Xenia Heilmann, Mattia Cerrato +2
Prior work evaluates code generation bias primarily through simple conditional statements, which represent only a narrow slice of real-world programming and reveal solely overt, ex…
Meenz bleibt Meenz, but Large Language Models Do Not Speak Its Dialect
Minh Duc Bui, Manuel Mager, Peter Herbert Kann +1
Meenzerisch, the dialect spoken in the German city of Mainz, is also the traditional language of the Mainz carnival, a yearly celebration well known throughout Germany. However, Me…
NALA_MAINZ at BLP-2025 Task 2: A Multi-agent Approach for Bangla Instruction to Python Code Generation
Hossain Shaikh Saadi, Faria Alam, Mario Sanz-Guerrero +3
This paper presents JGU Mainz's winning system for the BLP-2025 Shared Task on Code Generation from Bangla Instructions. We propose a multi-agent-based pipeline. First, a code-gene…