most citedRethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards

1 citations · 1 across the 16 of their papers we have counts for

collaborators

16 papers

cs.LG2026

Hyperparameter Scaling Laws Across MoE Sparsity

Changxin Tian, Kunlong Chen, Jia Liu +3

Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyperparameter transfer challengin…

cs.LG2026

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

Zihan Liu, Ruiheng Zheng, Shaobo Zhang +4

We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR).…

cs.CL2026

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Xinyu Tang, Qianggang Cao, Yurou Liu +13

Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoni…

cs.AI2026

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

Qian Zhao, Kunlong Chen, Changxin Tian +9

FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hardware paths and recipes, including NVIDIA Blackwell/Rubin-class…

cs.CL2026

PowLU: An Activation Function for Stable Pre-Training of LLMs

Peijie Jiang, Yuqi Feng, Cunyin Peng +5

In contemporary large language models (LLMs), the swish-gated linear unit (SwiGLU) activation function is widely adopted to regulate the information flow and introduce non-linearit…

cs.LG2026

Concordia: Self-Improving Synthetic Tables for Federated LLMs

Jimin Huang, Duanyu Feng, Nuo Chen +8

Federated learning (FL) enables training large language models (LLMs) without sharing raw data, but adapting LLMs under strict data isolation and non-IID client distributions remai…