5 papers
Cutting LLM Evaluation Costs with SySRs: A Bandit Algorithm that Provably Exploits Model Similarity
Zifan Lyu, Chahine Nejma, Tobias Wegel +2
Large Language Models are typically benchmarked by evaluating every model on every test query. For practitioners seeking the best model to deploy, this is often wasteful: if a mode…
Hedging on the Frontier: Learning New Tasks with Few Samples
Tobias Wegel, Federico Di Gennaro, Geelon So +1
When a learner faces a new task with few samples, it must leverage any available side information. In practice, this often comes in the form of model evaluations on related tasks i…
Time-sensitive anytime-valid testing
Eugenio Clerico, Tobias Wegel, Iskander Azangulov +1
Anytime-valid tests allow evidence to be checked during data collection: one can either continue testing or stop and reject the null while still controlling type-I error. Yet, in m…
On the sample complexity of semi-supervised multi-objective learning
Tobias Wegel, Geelon So, Junhyung Park +1
In multi-objective learning (MOL), several possibly competing prediction tasks must be solved jointly by a single model. Achieving good trade-offs may require a model class $\mathc…
Learning Pareto manifolds in high dimensions: How can regularization help?
Tobias Wegel, Filip KovaÄeviÄ, Alexandru Å¢ifrea +1
Simultaneously addressing multiple objectives is becoming increasingly important in modern machine learning. At the same time, data is often high-dimensional and costly to label. F…