Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Ronald Skorobogat, Ameya Prabhu, Matthias Bethge
Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to popular reasoning and knowledge…
cs.CL2025
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
Lukas Thede, Karsten Roth, Matthias Bethge +2
Keeping large language models factually up-to-date is crucial for deployment, yet costly retraining remains a challenge. Knowledge editing offers a promising alternative, but metho…
cs.CL2024
CiteME: Can Language Models Accurately Cite Scientific Claims?
Ori Press, Andreas Hochlehnert, Ameya Prabhu +3
Thousands of new scientific papers are published each month. Such information overload complicates researcher efforts to stay current with the state-of-the-art as well as to verify…