2 papers
cs.CL2026
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
Katelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi +5
Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-…
cs.SI2026
Grok in the Wild: Characterizing the Roles and Uses of Large Language Models on Social Media
Katelyn Xiaoying Mei, Robert Wolfe, Nicholas Weber +1
xAI's large language model, Grok, is called by millions of people each week on the social media platform X. Prior work characterizing how large language models are used has focused…