2 papers
cs.LG2026
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
Alexis Limozin, Eduard Durech, Torsten Hoefler +2
Recent mixed-policy optimization methods for LLM reasoning that interleave or blend supervised and reinforcement learning signals report improvements over the standard SFT-then-RL…
cs.LG2025
DISCO: A Browser-Based Privacy-Preserving Framework for Distributed Collaborative Learning
Julien T. T. Vignoud, Valérian Rousset, Hugo El Guedj +28
Data is often impractical to share for a range of well considered reasons, such as concerns over privacy, intellectual property, and legal constraints. This not only fragments the…