2 papers
cs.LG2026
Semi-Supervised Hypothesis Testing by Betting on Predictions
Yaniv Tenzer, Elad Tolochinsky, Yaniv Romano
We introduce a testing-by-betting framework that leverages predictions on unlabeled data to enhance the power of sequential hypothesis testing. Given limited samples from the joint…
cs.LG2026
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
Elad Tolochinsky, Yaniv Tenzer, Yaniv Romano
Selecting the best large language model (LLM) for a fixed benchmark is often expensive, since exhaustive evaluation requires running every model on every example. Multi-armed bandi…