2 papers
cs.LG2026
How Far Can Unsupervised RLVR Scale LLM Training?
Bingxiang He, Yuxin Zuo, Zeyuan Liu +18
Unsupervised reinforcement learning with verifiable rewards (URLVR) offers a pathway to scale LLM training beyond the supervision bottleneck by deriving rewards without ground trut…
math.ST2026
Maximal Ancillarity, Semiparametric Efficiency, and the Elimination of Nuisances
Marc Hallin, Bas J. M. Werker, Bo Zhou
Restricting statistical experiments via nuisance-ancillary -fields yields nuisance-free experiments. However, a moot point with ancillarity is that maximal ancillary -field…