4 papers
Quantifying Ranking Uncertainty in LLM Benchmarks
Bitya Neuhof, Yuval Benjamini
Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks. Rank confidence intervals were recently introduced as a method…
Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation
Bitya Neuhof, Yuval Benjamini
Pretrained models are often evaluated on multi-task leaderboards to measure their applicability in diverse contexts. However, current methods for aggregating performance across tas…
CardiCat: a Variational Autoencoder for High-Cardinality Tabular Data
Lee Carlin, Yuval Benjamini
High-cardinality categorical features are a common characteristic of mixed-type tabular datasets. Existing generative model architectures struggle to learn the complexities of such…
Class Distribution Shifts in Zero-Shot Learning: Learning Robust Representations
Yuli Slavutsky, Yuval Benjamini
Zero-shot learning methods typically assume that the new, unseen classes encountered during deployment come from the same distribution as the the classes in the training set. Howev…