2 citations · 2 across the 2 of their papers we have counts for
5 papers
Stochastic activations
Maria Lomeli, Matthijs Douze, Gergely Szilvasy +7
We introduce stochastic activations. This novel strategy randomly selects between several non-linear functions in the feed-forward layer of a large language model. In particular, w…
RTTC: Reward-Guided Collaborative Test-Time Compute
J. Pablo Muñoz, Jinjie Yuan
Test-Time Compute (TTC) has emerged as a powerful paradigm for enhancing the performance of Large Language Models (LLMs) at inference, leveraging strategies such as Test-Time Train…
Assay2Mol: large language model-based drug design using BioAssay context
Yifan Deng, Spencer S. Ericksen, Anthony Gitter
Scientific databases aggregate vast amounts of quantitative data alongside descriptive text. In biochemistry, molecule screening assays evaluate candidate molecules' functional res…
Inference-time sparse attention with asymmetric indexing
Pierre-Emmanuel Mazaré, Gergely Szilvasy, Maria Lomeli +4
Self-attention in transformer models is an incremental associative memory that maps key vectors to value vectors. One way to speed up self-attention is to employ GPU-compatible vec…
Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Aaditya K. Singh, Muhammed Yusuf Kocyigit, Andrew Poulton +4
Hampering the interpretation of benchmark scores, evaluation data contamination has become a growing concern in the evaluation of LLMs, and an active area of research studies its e…