2 papers
cs.LG2026
F-TIS: Harnessing Diverse Models in Collaborative GRPO
Nikolay Blagoev, OÄuzhan Ersoy, Wendelin Boehmer +1
Reinforcement learning methods such as GRPO have seen great popularity in LLM post-training. In GRPO, models produce completions to a set of prompts, which are rewarded, and the po…
cs.LG2026
Epistemic Monte Carlo Tree Search
Yaniv Oren, Viliam Vadocz, Matthijs T. J. Spaan +1
The AlphaZero/MuZero (A/MZ) family of algorithms has achieved remarkable success across various challenging domains by integrating Monte Carlo Tree Search (MCTS) with learned model…