1 paper
Teresa Yeo, Myeongho Jeon, Dulaj Weerakoon +4
We present a method for generating training data for reinforcement learning with verifiable rewards to improve small open-weights language models on mathematical tasks. Existing da…