4 papers
Not only where, But when: Temporal Scheduling for RLVR
Jinghao Zhang, Ruilin Li, Feng Zhao +1
Reinforcement learning with verifiable rewards (RLVR) has become a core technique for post-training of Large Language Models (LLMs). While policy optimization is driven by all samp…
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
MiroMind Team, Song Bai, Lidong Bing +52
We present MiroThinker v1.0, an open-source research agent designed to advance tool-augmented reasoning and information-seeking capabilities. Unlike previous agents that only scale…
Enhancing Pretrained Model-based Continual Representation Learning via Guided Random Projection
Ruilin Li, Heming Zou, Xiufeng Yan +4
Recent paradigms in Random Projection Layer (RPL)-based continual representation learning have demonstrated superior performance when building upon a pre-trained model (PTM). These…
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
Xingxuan Li, Yao Xiao, Dianwen Ng +15
Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains,…