5 papers
When Do Larger Batches Help Scale LLM Reinforcement Learning?
Ziniu Li, Jinbo Wang, Guanhua Huang +3
Larger batches reduce the variance of stochastic gradients per update and are therefore often expected to accelerate training. Yet whether this statistical benefit translates into…
Complementary Quantum Correlations Are Universal for Qubits
Jinbo Wang, Qihang Wang, Kun Chen
Extracting total correlations from a quantum system usually requires reconstructing its state, whereas many experiments access only a few measurement settings. A possible shortcut…
When Complementary Measurements Count the Same Classical Bit Twice: Counterexamples to CQC, ECQC, and Complementarity-Based Certification
Jinbo Wang, Qihang Wang, Kun Chen
Mutually unbiased measurements are commonly expected to expose independent facets of a quantum state: a correlation that is classical in one basis should disappear in a complementa…
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
Jinbo Wang, Binghui Li, Zhanpeng Zhou +5
Batch size scheduling (BSS) plays a critical role in large-scale deep learning training, influencing both optimization dynamics and computational efficiency. Yet, its theoretical f…
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
Jinbo Wang, Mingze Wang, Zhanpeng Zhou +3
Transformers consist of diverse building blocks, such as embedding layers, normalization layers, self-attention mechanisms, and point-wise feedforward networks. Thus, understanding…