4 papers
GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity
Yong Yi Bay, Kathleen A. Yearick
Three of the most popular methods for training language models to reason look like three different tricks. They are not. All three adjust a single number: standard deviation, refle…
When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling
Yong Yi Bay, Kathleen A. Yearick
People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning systems answer a hard question by sampling it many times (test-time…
No 3D Matrices: A Unified Tensor-Product View of Matrix-Free Cartesian PDE Solvers
Yong Yi Bay, Kathleen A. Yearick
Every Cartesian three-dimensional PDE solver hides a structural secret that production CFD codes have used for half a century and that graduate-level textbooks rarely state plainly…
Solve for the Hyperparameter, Skip the Search: Kolmogorov-Optimal Scaling Laws for Spline Regression
Yong Yi Bay, Kathleen A. Yearick
Hyperparameter tuning almost always means search: fit the model at every value on a grid, score each by cross-validation, and keep the winner. For spline regression that search is…