1 citations · 1 across the 13 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
JTok: On Token Embedding as another Axis of Scaling Law via Joint Token Self-modulation
Yebin Yang, Huaijin Wu, Fu Guo +5
LLMs have traditionally scaled along dense dimensions, where performance is coupled with near-linear increases in computational cost. While MoE decouples capacity from compute, it…
cs.LG2025
AdaMuon: Adaptive Muon Optimizer
Chongjie Si, Debing Zhang, Wei Shen
We propose AdaMuon, a novel optimizer that combines element-wise adaptivity with orthogonal updates for large-scale neural network training. AdaMuon incorporates two tightly couple…
cs.LG2025★ 1 cited
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
Haotian Xu, Xing Wu, Weinong Wang +11
Can scaling transform reasoning? In this work, we explore the untapped potential of scaling Long Chain-of-Thought (Long-CoT) data to 1000k samples, pioneering the development of a…