4 papers
MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
Jiacheng Chen, Xinyu Zhang, Shunkai Zhang +20
We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabili…
Operator-Guided Invariance Learning for Continuous Reinforcement Learning
Zuyuan Zhang, Fei Xu Yu, Tian Lan
Reinforcement learning (RL) with continuous time and state/action spaces is often data-intensive and brittle under nuisance variability and shift, motivating methods that exploit v…
ACDZero: MCTS Agent for Mastering Automated Cyber Defense
Yu Li, Sizhe Tang, Rongqian Chen +5
Automated cyber defense (ACD) seeks to protect computer networks with minimal or no human intervention, reacting to intrusions by taking corrective actions such as isolating hosts,…
Optimizing Prompt Sequences using Monte Carlo Tree Search for LLM-Based Optimization
Fei Xu Yu, Gina Adam, Nathaniel D. Bastian +1
Large language models (LLMs) have demonstrated remarkable capabilities in code generation and structured reasoning; however, their performance often degrades on complex tasks that…