4 papers · 1 filter
Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
Sherzod Hakimov, Yerkezhan Abdullayeva, Kushal Koshti +4
While the situation has improved for text-only models, it again seems to be the case currently that multimodal (text and image) models develop faster than ways to evaluate them. In…
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
Anne Beyer, Kranti Chalamalasetti, Sherzod Hakimov +3
It has been established in recent work that Large Language Models (LLMs) can be prompted to "self-play" conversational games that probe certain capabilities (general instruction fo…
Neural Conversation Models and How to Rein Them in: A Survey of Failures and Fixes
Fabian Galetzka, Anne Beyer, David Schlangen
Recent conditional language models are able to continue any kind of text source in an often seemingly fluent way. This fact encouraged research in the area of open-domain conversat…
Is Incoherence Surprising? Targeted Evaluation of Coherence Prediction from Language Models
Anne Beyer, Sharid Loáiciga, David Schlangen
Coherent discourse is distinguished from a mere collection of utterances by the satisfaction of a diverse set of constraints, for example choice of expression, logical relation bet…