Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Benchmarking LLMs' Judgments with No Gold Standard
Shengwei Xu, Yuxuan Lu, Grant Schoenebeck +1
We introduce the GEM (Generative Estimator for Mutual Information), an evaluation metric for assessing language generation by Large Language Models (LLMs), particularly in generati…
cs.CL2024
Eliciting Informative Text Evaluations with Large Language Models
Yuxuan Lu, Shengwei Xu, Yichi Zhang +2
Peer prediction mechanisms motivate high-quality feedback with provable guarantees. However, current methods only apply to rather simple reports, like multiple-choice or scalar num…