5 citations · 5 across the 2 of their papers we have counts for
4 papers
Queer NLP: A Critical Survey on Literature Gaps, Biases and Trends
Sabine Weber, Angelina Wang, Ankush Gupta +16
Natural language processing (NLP) technologies are rapidly reshaping how language is created, processed, and interpreted by humans. With current and potential applications in hirin…
Lessons from the Trenches on Reproducible Evaluation of Language Models
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika +27
Reliable evaluation of language models (LMs) remains an open challenge. Re- searchers and engineers face methodological issues such as the sensitivity of models to evaluation setup…
Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks
Wilson Y. Lee
Why do language agents fail on tasks they are capable of solving? We argue that many such failures are reliability failures caused by stochastic drift from a task's latent solution…
How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation
Wilson Y. Lee
Human preference evaluations are widely used to compare generative models, yet it remains unclear how many judgments are required to reliably detect small improvements. We show tha…