9 papers
DEEPRUBRIC: Evidence-Tree Rubric Supervision for Efficient Reinforcement Learning of Deep Research Agents
Minghang Zhu, Chuyang Wei, Junhao Xu +3
Deep research agents synthesize long-form reports by searching and reasoning over retrieved evidence. Reinforcement learning with rubric-based rewards improves these agents by opti…
One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents
Zhaoxi Zhang, Yitong Duan, Yanzhi Zhang +9
Locating files and functions requiring modification in large software repositories is challenging due to their scale and structural complexity. Existing LLM-based methods typically…
FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards
Zhixin Han, Yanzhi Zhang, Chuyang Wei +11
Live future prediction refers to the task of making predictions about real-world events before they unfold. This task is increasingly studied using large language model-based agent…
Harnessing Pre-Resolution Signals for Future Prediction Agents
Chuyang Wei, Maohang Gao, Zhixin Han +12
Many high-stakes decisions depend on forecasts made before outcomes are known. In this future prediction setting, the central challenge is that public evidence evolves over time, w…
Can a Lightweight Automated AI Pipeline Solve Research-Level Mathematical Problems?
Lve Meng, Weilong Zhao, Yanzhi Zhang +2
Large language models (LLMs) have recently achieved remarkable success in generating rigorous mathematical proofs, with "AI for Math" emerging as a vibrant field of research (Ju et…
Population-Evolve: a Parallel Sampling and Evolutionary Method for LLM Math Reasoning
Yanzhi Zhang, Yitong Duan, Zhaoxi Zhang +2
Test-time scaling has emerged as a promising direction for enhancing the reasoning capabilities of Large Language Models in last few years. In this work, we propose Population-Evol…