◍wovepaper
PapersResearchersInstitutions
Sign in
researcher

Martin Forell

1 papers

No researched profile yet.

papers

Publications (1)

cs.CL2026

(Towards) Scalable Reliable Automated Evaluation with Large Language Models

Bertil Braun, Martin Forell

The paper presents a scalable framework for automatically evaluating large language model outputs using pairwise comparisons and an Elo rating system, achieving rankings that align…

#automated evaluation#large language models#pairwise comparison#elo rating
◍wovepaper

A living map of arXiv — papers, researchers, institutions.

Explore
  • Papers
  • Researchers
  • Institutions
Account
  • Sign in
  • For you
  • Library
  • Chat
Data
  • arXiv.org
  • Latest RSS
Metadata from arXiv.org · Not affiliated with arXiv