4 papers
Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics
Shengwei Xu, Yuxuan Lu, Yifan Wu +2
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response.…
Stochastically Dominant Peer Prediction
Yichi Zhang, Shengwei Xu, David Pennock +1
Eliciting reliable human feedback is essential for many machine learning tasks, such as learning from noisy labels and aligning AI systems with human preferences. Peer prediction m…
Ad Insertion in LLM-Generated Responses
Shengwei Xu, Zhaohua Chen, Xiaotie Deng +2
Sustainable monetization of large language models (LLMs) remains a critical open challenge. Traditional search advertising, which relies on static keywords, fails to capture the fl…
Benchmarking LLMs' Judgments with No Gold Standard
Shengwei Xu, Yuxuan Lu, Grant Schoenebeck +1
We introduce the GEM (Generative Estimator for Mutual Information), an evaluation metric for assessing language generation by Large Language Models (LLMs), particularly in generati…