Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation
Bitya Neuhof, Yuval Benjamini
Pretrained models are often evaluated on multi-task leaderboards to measure their applicability in diverse contexts. However, current methods for aggregating performance across tas…
stat.ML2024
Confident Feature Ranking
Bitya Neuhof, Yuval Benjamini
Machine learning models are widely applied in various fields. Stakeholders often use post-hoc feature importance methods to better understand the input features' contribution to th…