3 papers
cs.LG2026
Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions
Lingkai Kong, Anagha Satish, Hezi Jiang +6
Reinforcement learning (RL) with combinatorial action spaces remains challenging because feasible action sets are exponentially large and governed by complex feasibility constraint…
cs.LG2025
Practical Efficiency of Muon for Pretraining
Essential AI, :, Ishaan Shah +22
We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon…
cs.CL2025
Rethinking Reflection in Pre-Training
Essential AI, :, Darsh J Shah +26
A language model's ability to reflect on its own reasoning provides a key advantage for solving complex problems. While most recent research has focused on how this ability develop…