collaborators

5 papers

cs.LG2026

Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards

Sarah Liaw, Benjamin Plaut

In high-stakes AI applications, even a single action can cause irreparable damage. However, nearly all of sequential decision-making theory assumes that all errors are recoverable…

cs.LG2026

Graphon Mean-Field Subsampling for Cooperative Heterogeneous Multi-Agent Reinforcement Learning

Emile Anand, Richard Hoffmann, Sarah Liaw +1

Coordinating large populations of interacting agents is a central challenge in multi-agent reinforcement learning (MARL), where the size of the joint state-action space scales expo…

cs.AI2026

FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale

Isabelle Lee, Sarah Liaw, Dani Yogatama

Reasoning in language models is difficult to evaluate: natural-language traces are unverifiable, symbolic datasets are too small, and most benchmarks conflate heuristics with infer…

cs.LG2025

Feel-Good Thompson Sampling for Contextual Bandits: a Markov Chain Monte Carlo Showdown

Emile Anand, Sarah Liaw

Thompson Sampling (TS) is widely used to address the exploration/exploitation tradeoff in contextual bandits, yet recent theory shows that it does not explore aggressively enough i…

cs.LG2025

Learning local neighborhoods of non-Gaussian graphical models: A measure transport approach

Sarah Liaw, Rebecca Morrison, Youssef Marzouk +1

Identifying the Markov properties or conditional independencies of a collection of random variables is a fundamental task in statistics for modeling and inference. Existing approac…