4 papers · 1 filter
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
Ikram Belmadani, Oumaima El Khettari, Pacôme Constant dit Beaufils +2
Automatic evaluation of medical open-ended question answering (OEQA) remains challenging due to the need for expert annotations. We evaluate whether large language models (LLMs) ca…
Coupling Local Context and Global Semantic Prototypes via a Hierarchical Architecture for Rhetorical Roles Labeling
Anas Belfathi, Nicolas Hernandez, Laura Monceaux +4
Rhetorical Role Labeling (RRL) identifies the functional role of each sentence in a document, a key task for discourse understanding in domains such as law and medicine. While hier…
Identifying Reliable Evaluation Metrics for Scientific Text Revision
Léane Jourdan, Florian Boudin, Richard Dufour +1
Evaluating text revision in scientific writing remains a challenge, as traditional metrics such as ROUGE and BERTScore primarily focus on similarity rather than capturing meaningfu…
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
Léane Jourdan, Nicolas Hernandez, Richard Dufour +2
Revision is a crucial step in scientific writing, where authors refine their work to improve clarity, structure, and academic quality. Existing approaches to automated writing assi…