activity
20242026
collaborators

6 papers

stat.ML2026

Efficient Evaluation of LLM Performance with Statistical Guarantees

Skyler Wu, Yash Nair, Emmanuel J. Candès

Exhaustively evaluating many large language models (LLMs) on a large suite of benchmarks is expensive. We cast benchmarking as finite-population inference and, under a fixed query…

stat.CO2026

Are Statistical Methods Obsolete in the Era of Deep Learning? A Study of ODE Inverse Problems

Skyler Wu, Shihao Yang, S. C. Kou

In the era of AI, neural networks have become increasingly popular for modeling, inference, and prediction, largely due to their potential for universal approximation. With the pro…

cs.LG2025

Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs

Skyler Wu, Aymen Echarghaoui

Fine-tuning is integral for aligning large language models (LLMs) with human preferences. Multiple-Reference Preference Optimization (MRPO) builds on Direct Preference Optimization…

stat.CO2025

Parallelizing MCMC Across the Sequence Length

David M. Zoltowski, Skyler Wu, Xavier Gonzalez +2

Markov chain Monte Carlo (MCMC) methods are foundational algorithms for Bayesian inference and probabilistic modeling. However, most MCMC algorithms are inherently sequential and t…

stat.ML2025

Missing Data Multiple Imputation for Tabular Q-Learning in Online RL

Kyla Chasalow, Skyler Wu, Susan Murphy

Missing data in online reinforcement learning (RL) poses challenges compared to missing data in standard tabular data or in offline policy learning. The need to impute and act at e…

cs.LG2024

Stabilizing Linear Passive-Aggressive Online Learning with Weighted Reservoir Sampling

Skyler Wu, Fred Lu, Edward Raff +1

Online learning methods, like the seminal Passive-Aggressive (PA) classifier, are still highly effective for high-dimensional streaming data, out-of-core processing, and other thro…