5 papers
Interpretability Can Be Actionable
Hadas Orgad, Fazl Barez, Tal Haklay +9
Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impa…
Rigorous Interpretation Is a Form of Evaluation
Isabelle Lee, Emmy Liu, Cathy Jiao +4
Current machine learning models are evaluated through behavioral snapshots, with benchmark accuracies, win rates and outcome-based metrics. Model explanations and evaluations, howe…
What do Language Models Learn and When? The Implicit Curriculum Hypothesis
Emmy Liu, Kaiser Sun, Millicent Li +4
Large language models (LLMs) can perform remarkably complex tasks, yet the fine-grained details of how these capabilities emerge during pretraining remain poorly understood. Scalin…
FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale
Isabelle Lee, Sarah Liaw, Dani Yogatama
Reasoning in language models is difficult to evaluate: natural-language traces are unverifiable, symbolic datasets are too small, and most benchmarks conflate heuristics with infer…
Evaluating Large Language Models for Fair and Reliable Organ Allocation
Brian Hyeongseok Kim, Hannah Murray, Isabelle Lee +4
Medical institutions are considering the use of LLMs in high-stakes clinical decision-making, such as organ allocation. In such sensitive use cases, evaluating fairness is imperati…