3 papers
cs.AI2025
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
Bo Liu, Leon Guertler, Simon Yu +9
Recent advances in reinforcement learning have shown that language models can develop sophisticated reasoning through training on tasks with verifiable rewards, but these approache…
cs.CL2024
STLM Engineering Report: Dropout
Dylan Hillier, Leon Guertler, Bobby Cheng +1
In this work we explore the relevance of dropout for modern language models, particularly in the context of models on the scale of <100M parameters. We explore it's relevance first…
cs.LG2022
How to train your draGAN: A task oriented solution to imbalanced classification
Leon O. Guertler, Andri Ashfahani, Anh Tuan Luu
The long-standing challenge of building effective classification models for small and imbalanced datasets has seen little improvement since the creation of the Synthetic Minority O…