3 papers
cs.LG2025
Alignment of large language models with constrained learning
Botong Zhang, Shuo Li, Ignacio Hounie +3
We study the problem of computing an optimal large language model (LLM) policy for the constrained alignment problem, where the goal is to maximize a primary reward objective while…
cs.CL2025
Evaluating the Diversity and Quality of LLM Generated Content
Alexander Shypula, Shuo Li, Botong Zhang +3
Recent work suggests that preference-tuning techniques -- such as Reinforcement Learning from Human Feedback (RLHF) methods like PPO and GRPO, as well as alternatives like DPO -- r…
cs.LG2024
Conformal Structured Prediction
Botong Zhang, Shuo Li, Osbert Bastani
Conformal prediction has recently emerged as a promising strategy for quantifying the uncertainty of a predictive model; these algorithms modify the model to output sets of labels…