2 papers
cs.AI2026
Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings
Mina Remeli, Moritz Hardt
Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues…
cs.CL2026
Limits to Predicting Online Speech Using Large Language Models
Mina Remeli, Moritz Hardt, Robert C. Williamson
Our paper studies the predictability of online speech -- that is, how well language models learn to model the distribution of user generated content on X (previously Twitter). We d…