5 papers
Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics
Shengwei Xu, Yuxuan Lu, Yifan Wu +2
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response.…
AnnotateThis: Analyzing a human-LLM system for annotating social media data with the concept of climate change mitigation pessimism
Zexuan Li, Derek Van Berkel, Ariel Hasell +3
Large language models (LLMs) are increasingly being integrated into research workflows. However, LLMs have been shown to struggle with difficult and nuanced concepts such as those…
Stochastically Dominant Peer Prediction
Yichi Zhang, Shengwei Xu, David Pennock +1
Eliciting reliable human feedback is essential for many machine learning tasks, such as learning from noisy labels and aligning AI systems with human preferences. Peer prediction m…
Ad Insertion in LLM-Generated Responses
Shengwei Xu, Zhaohua Chen, Xiaotie Deng +2
Sustainable monetization of large language models (LLMs) remains a critical open challenge. Traditional search advertising, which relies on static keywords, fails to capture the fl…
Benchmarking LLMs' Judgments with No Gold Standard
Shengwei Xu, Yuxuan Lu, Grant Schoenebeck +1
We introduce the GEM (Generative Estimator for Mutual Information), an evaluation metric for assessing language generation by Large Language Models (LLMs), particularly in generati…