1 paper
Naifan Zhang, Ruihan Sun, Ruixi Su +9
The LLM field has spent a year perfecting RL for tasks machines already excel at, math, code, and deterministic reasoning, while completely sidestepping the domain that actually de…