2 papers
cs.AI2026
Aligning Language Model Benchmarks with Pairwise Preferences
Marco Gutierrez, Xinyi Leng, Hannah Cyberey +3
Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find that benchmarks often fail to predict real…
cs.CL2024
Narrative Analysis of True Crime Podcasts With Knowledge Graph-Augmented Large Language Models
Xinyi Leng, Jason Liang, Jack Mauro +8
Narrative data spans all disciplines and provides a coherent model of the world to the reader or viewer. Recent advancement in machine learning and Large Language Models (LLMs) hav…