Showing 2025 · cs.LGShow all
2 papers · 2 filters
cs.LG2025
LLM generation novelty through the lens of semantic similarity
Philipp Davydov, Ameya Prabhu, Matthias Bethge +2
Generation novelty is a key indicator of an LLM's ability to generalize, yet measuring it against full pretraining corpora is computationally challenging. Existing evaluations ofte…
cs.LG2025
DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
Alexander Rubinstein, Benjamin Raible, Martin Gubri +1
Evaluating modern machine learning models has become prohibitively expensive. Benchmarks such as LMMs-Eval and HELM demand thousands of GPU hours per model. Costly evaluation reduc…