automated evaluation 1domain-agnostic assessment 1elo rating 1large language models 1pairwise comparison 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
Bertil Braun, Martin Forell
The paper presents a scalable framework for automatically evaluating large language model outputs using pairwise comparisons and an Elo rating system, achieving rankings that align…
cs.LG2026
A Graph-Based Control Interface for Traffic Signals on Heterogeneous Road Networks
Bertil Braun
We present a traffic-signal control interface in which a shared graph neural network assigns scores to individual traffic movements. Each junction converts these scores into its ow…