2 papers
cs.AI2026
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence
Xinquan Chen, Zhenyun Yin, Shan He +38
As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool use, and real environment intera…
cs.AI2025
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
Bingning Huang, Tu Nguyen, Matthieu Zimmer
Recent advances in reasoning with large language models (LLMs) have shown the effectiveness of Monte Carlo Tree Search (MCTS) for generating high quality intermediate trajectories,…