7 papers
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
Jasper Dekoninck, Nikola JovanoviÄ, Tim Gehrunger +4
Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating progress: they are often na…
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
Mislav BalunoviÄ, Jasper Dekoninck, Nikola JovanoviÄ +2
While Large Language Models (LLMs) demonstrate impressive performance in mathematics, existing math benchmarks come with significant limitations. Many focus on problems with fixed…
BaxBench: Can LLMs Generate Correct and Secure Backends?
Mark Vero, Niels Mündler, Victor Chibotaru +5
Automatic program generation has long been a fundamental challenge in computer science. Recent benchmarks have shown that large language models (LLMs) can effectively generate code…
Discovering Spoofing Attempts on Language Model Watermarks
Thibaud Gloaguen, Nikola JovanoviÄ, Robin Staab +1
LLM watermarks stand out as a promising way to attribute ownership of LLM-generated text. One threat to watermark credibility comes from spoofing attacks, where an unauthorized thi…
Ward: Provable RAG Dataset Inference via LLM Watermarks
Nikola JovanoviÄ, Robin Staab, Maximilian Baader +1
RAG enables LLMs to easily incorporate external data, raising concerns for data owners regarding unauthorized usage of their content. The challenge of detecting such unauthorized u…
Towards Watermarking of Open-Source LLMs
Thibaud Gloaguen, Nikola JovanoviÄ, Robin Staab +1
While watermarks for closed LLMs have matured and have been included in large-scale deployments, these methods are not applicable to open-source models, which allow users full cont…