Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Mapping Similarity Spaces across Embedding Models with Synthetic Query Probing
Marcin Rozmus, Peter van der Putten
Retrieval-Augmented Generation systems rely on similarity scores to retrieve relevant content, yet scores are not directly comparable across embedding models due to differing geome…
cs.CL2026
Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
Ian B. de Haan, Peter van der Putten, Max van Duijn
Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the underlying capabilities. At…
cs.CL2025
Can we Evaluate RAGs with Synthetic Data?
Jonas van Elburg, Peter van der Putten, Maarten Marx
We investigate whether synthetic question-answer (QA) data generated by large language models (LLMs) can serve as an effective proxy for human-labeled benchmarks when the latter is…