4 papers · 1 filter
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
David Schlangen, Sherzod Hakimov, Chalamalasetti Kranti +2
There are currently two main paradigms for evaluating large language models (LLMs), reference-based evaluation and preference-based evaluation. The first, carried over from the eva…
Playpen: An Environment for Exploring Learning Through Conversational Interaction
Nicola Horst, Davide Mazzaccara, Antonia Schmidt +13
Interaction between learner and feedback-giver has come into focus recently for post-training of Large Language Models (LLMs), through the use of reward models that judge the appro…
The Unreasonable Ineffectiveness of Nucleus Sampling on Mitigating Text Memorization
Luka Borec, Philipp Sadler, David Schlangen
This work analyses the text memorization behavior of large language models (LLMs) when subjected to nucleus sampling. Stochastic decoding methods like nucleus sampling are typicall…
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
Anne Beyer, Kranti Chalamalasetti, Sherzod Hakimov +3
It has been established in recent work that Large Language Models (LLMs) can be prompted to "self-play" conversational games that probe certain capabilities (general instruction fo…