most citedEvaluation data contamination in LLMs: how do we measure it and (when) does it matter?

2 citations · 2 across the 2 of their papers we have counts for

collaborators

5 papers

cs.LG2025

Stochastic activations

Maria Lomeli, Matthijs Douze, Gergely Szilvasy +7

We introduce stochastic activations. This novel strategy randomly selects between several non-linear functions in the feed-forward layer of a large language model. In particular, w…

cs.CL2025

RTTC: Reward-Guided Collaborative Test-Time Compute

J. Pablo Muñoz, Jinjie Yuan

Test-Time Compute (TTC) has emerged as a powerful paradigm for enhancing the performance of Large Language Models (LLMs) at inference, leveraging strategies such as Test-Time Train…

cs.LG2025

Assay2Mol: large language model-based drug design using BioAssay context

Yifan Deng, Spencer S. Ericksen, Anthony Gitter

Scientific databases aggregate vast amounts of quantitative data alongside descriptive text. In biochemistry, molecule screening assays evaluate candidate molecules' functional res…

cs.CL2025

Inference-time sparse attention with asymmetric indexing

Pierre-Emmanuel Mazaré, Gergely Szilvasy, Maria Lomeli +4

Self-attention in transformer models is an incremental associative memory that maps key vectors to value vectors. One way to speed up self-attention is to employ GPU-compatible vec…

cs.CL20242 cited

Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Aaditya K. Singh, Muhammed Yusuf Kocyigit, Andrew Poulton +4

Hampering the interpretation of benchmark scores, evaluation data contamination has become a growing concern in the evaluation of LLMs, and an active area of research studies its e…