Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Conditional Factuality Controlled LLMs with Generalization Certificates via Conformal Sampling
Kai Ye, Qingtao Pan, Shuo Li
Large language models (LLMs) need reliable test-time control of hallucinations. Existing conformal methods for LLMs typically provide only \emph{marginal} guarantees and rely on a…
cs.LG2026
Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-Statistic
Hongyi Zhou, Kai Ye, Erhan Xu +4
Group relative policy optimization (GRPO), a core methodological component of DeepSeekMath and DeepSeek-R1, has emerged as a cornerstone for scaling reasoning capabilities of large…
cs.LG2025
Doubly Robust Alignment for Large Language Models
Erhan Xu, Kai Ye, Hongyi Zhou +3
This paper studies reinforcement learning from human feedback (RLHF) for aligning large language models with human preferences. While RLHF has demonstrated promising results, many…