1 paper
Dushyant Rajput, Nirdesh Chauhan, Siddharth Kosaraju
We scale our conventional sub-150M pretraining recipe from 53.5M to 109.7M parameters, holding the method fixed (Qwen3-style decoder with grouped-query attention, RoPE, SwiGLU, RMS…