4 papers
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach +6
We introduce BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in German law. The dataset combines 596 exam…
BenGER Platform: A Collaborative Web Platform for End-to-End Benchmarking of German Legal Tasks
Sebastian Nagl, Matthias Grabmair
Evaluating large language models (LLMs) for legal reasoning requires workflows that span task design, expert annotation, model execution, and metric-based evaluation. In practice,…
Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization
Mohamed Hesham Elganayni, Runsheng Chen, Sebastian Nagl +1
This work explores the role of prompt design and judge selection in LLM-as-a-Judge evaluations of free text legal question answering. We examine whether automatic task prompt optim…
CourtPressGER: A German Court Decision to Press Release Summarization Dataset
Sebastian Nagl, Mohamed Elganayni, Melanie Pospisil +1
Official court press releases from Germany's highest courts present and explain judicial rulings to the public, as well as to expert audiences. Prior NLP efforts emphasize technica…