8 papers · 1 filter
Example-Guided Prompting for Document-Level Text Simplification
Marina Litvak, Ariel Perstin, Ilan Shtilman +1
Document-level text simplification requires large language models (LLMs) to rewrite complex documents while preserving meaning, readability, and discourse coherence. Although promp…
Topic-to-Timestamp Alignment by Constrained Evidence Selection
Zeynep Yılbırt, Marina Litvak, Michael Färber
Meeting archives are difficult to search when users remember what was discussed but not when. We study topic-to-timestamp alignment: given a natural-language topic and a timestampe…
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
Klaudia-Doris Thellmann, Bernhard Stadler, Michael Färber +1
Machine-translated benchmarks are widely used to assess the multilingual capabilities of large language models (LLMs), yet translation errors in these benchmarks remain underexplor…
Tracing Relational Knowledge Recall in Large Language Models
Nicholas PopoviÄ, Michael Färber
We study how large language models recall relational knowledge during text generation, with a focus on identifying latent representations suitable for relation classification via l…
Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite
Klaudia Thellmann, Bernhard Stadler, Michael Färber
Machine-translated benchmark datasets reduce costs and offer scale, but noise, loss of structure, and uneven quality weaken confidence. What matters is not merely whether we can tr…
Benchmarking Uncertainty Calibration in Large Language Model Long-Form Question Answering
Philip Müller, Nicholas PopoviÄ, Michael Färber +1
Large Language Models (LLMs) are commonly used in Question Answering (QA) settings, increasingly in the natural sciences if not science at large. Reliable Uncertainty Quantificatio…