7 papers
Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
Yanwei Ren, Haotian Zhang, Likang Xiao +5
Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete or misleading behaviors. Howe…
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
Chenghao Li, Fusheng Hao, Xikai Zhang +5
Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they remain limited in long-horizo…
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
Yanwei Ren, Haotian Zhang, Likang Xiao +6
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for enhancing the complex reasoning capabilities of Large Reasoning Models. However, standa…
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
Haotian Zhang, Liu Liu, Baosheng Yu +5
Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time sca…
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
Yanwei Ren, Liu Liu, Baosheng Yu +2
Optimizing instructions for large language models (LLMs) is critical for harnessing their full potential in complex and diverse tasks. However, relying solely on white-box approach…
LARGO: Low-Rank Regulated Gradient Projection for Robust Parameter Efficient Fine-Tuning
Haotian Zhang, Liu Liu, Baosheng Yu +3
The advent of parameter-efficient fine-tuning methods has significantly reduced the computational burden of adapting large-scale pretrained models to diverse downstream tasks. Howe…