7 papers
Mapping Similarity Spaces across Embedding Models with Synthetic Query Probing
Marcin Rozmus, Peter van der Putten
Retrieval-Augmented Generation systems rely on similarity scores to retrieve relevant content, yet scores are not directly comparable across embedding models due to differing geome…
Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
Ian B. de Haan, Peter van der Putten, Max van Duijn
Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the underlying capabilities. At…
Rashomon Alignment
Moisés Santos, Peter van der Putten, Bernhard Pfahringer +1
We propose Rashomon Alignment (RA), a new measure to assess functional similarity between two models. Existing functional similarity measures are distributional, quantifying differ…
Reasoning Promotes Robustness in Theory of Mind Tasks
Ian B. de Haan, Peter van der Putten, Max van Duijn
Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and true performance of the underlying capabili…
Mirror Mode in Fire Emblem: Beating Players at their own Game with Imitation and Reinforcement Learning
Yanna Elizabeth Smid, Peter van der Putten, Aske Plaat
Enemy strategies in turn-based games should be surprising and unpredictable. This study introduces Mirror Mode, a new game mode where the enemy AI mimics the personal strategy of a…
Agentic Large Language Models, a survey
Aske Plaat, Max van Duijn, Niki van Stein +3
Background: There is great interest in agentic LLMs, large language models that act as agents. Objectives: We review the growing body of work in this area and provide a research ag…