collaborators

6 papers

cs.LG2026

Interpretability Can Be Actionable

Hadas Orgad, Fazl Barez, Tal Haklay +9

Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impa…

cs.CY2026

Rigorous Interpretation Is a Form of Evaluation

Isabelle Lee, Emmy Liu, Cathy Jiao +4

Current machine learning models are evaluated through behavioral snapshots, with benchmark accuracies, win rates and outcome-based metrics. Model explanations and evaluations, howe…

cs.CL2026

What do Language Models Learn and When? The Implicit Curriculum Hypothesis

Emmy Liu, Kaiser Sun, Millicent Li +4

Large language models (LLMs) can perform remarkably complex tasks, yet the fine-grained details of how these capabilities emerge during pretraining remain poorly understood. Scalin…

cs.AI2026

FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale

Isabelle Lee, Sarah Liaw, Dani Yogatama

Reasoning in language models is difficult to evaluate: natural-language traces are unverifiable, symbolic datasets are too small, and most benchmarks conflate heuristics with infer…

cs.LG2026

Evaluating Large Language Models for Fair and Reliable Organ Allocation

Brian Hyeongseok Kim, Hannah Murray, Isabelle Lee +4

Medical institutions are considering the use of LLMs in high-stakes clinical decision-making, such as organ allocation. In such sensitive use cases, evaluating fairness is imperati…

cs.CL2024

Self-Contradictory Reasoning Evaluation and Detection

Ziyi Liu, Soumya Sanyal, Isabelle Lee +4

In a plethora of recent work, large language models (LLMs) demonstrated impressive reasoning ability, but many proposed downstream reasoning tasks only focus on final answers. Two…