collaborators

5 papers

cs.AI2026

Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

Shengwei Xu, Yuxuan Lu, Yifan Wu +2

Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response.…

cs.CY2026

AnnotateThis: Analyzing a human-LLM system for annotating social media data with the concept of climate change mitigation pessimism

Zexuan Li, Derek Van Berkel, Ariel Hasell +3

Large language models (LLMs) are increasingly being integrated into research workflows. However, LLMs have been shown to struggle with difficult and nuanced concepts such as those…

cs.GT2026

Stochastically Dominant Peer Prediction

Yichi Zhang, Shengwei Xu, David Pennock +1

Eliciting reliable human feedback is essential for many machine learning tasks, such as learning from noisy labels and aligning AI systems with human preferences. Peer prediction m…

cs.GT2026

Ad Insertion in LLM-Generated Responses

Shengwei Xu, Zhaohua Chen, Xiaotie Deng +2

Sustainable monetization of large language models (LLMs) remains a critical open challenge. Traditional search advertising, which relies on static keywords, fails to capture the fl…

cs.CL2025

Benchmarking LLMs' Judgments with No Gold Standard

Shengwei Xu, Yuxuan Lu, Grant Schoenebeck +1

We introduce the GEM (Generative Estimator for Mutual Information), an evaluation metric for assessing language generation by Large Language Models (LLMs), particularly in generati…