3 papers
cs.LG2025
Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs
Skyler Wu, Aymen Echarghaoui
Fine-tuning is integral for aligning large language models (LLMs) with human preferences. Multiple-Reference Preference Optimization (MRPO) builds on Direct Preference Optimization…
stat.ML2025
Missing Data Multiple Imputation for Tabular Q-Learning in Online RL
Kyla Chasalow, Skyler Wu, Susan Murphy
Missing data in online reinforcement learning (RL) poses challenges compared to missing data in standard tabular data or in offline policy learning. The need to impute and act at e…
stat.CO2025
Parallelizing MCMC Across the Sequence Length
David M. Zoltowski, Skyler Wu, Xavier Gonzalez +2
Markov chain Monte Carlo (MCMC) methods are foundational algorithms for Bayesian inference and probabilistic modeling. However, most MCMC algorithms are inherently sequential and t…