Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Nonparametric LLM Evaluation from Preference Data
Dennis Frauen, Athiya Deviyani, Mihaela van der Schaar +1
Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards. However, many existing approaches either rely on re…
cs.LG2026
Causal methods for LLM development and evaluation
Dennis Frauen, Marie Brockschmidt, Konstantin Hess +10
Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and evaluation pipelines. Here,…