collaborators

5 papers

cs.AI2026

Targeted Exploration via Unified Entropy Control for Reinforcement Learning

Chen Wang, Lai Wei, Yanzhi Zhang +5

Recent advances in reinforcement learning (RL) have improved the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs). However, the widely used…

cs.AI2026

Can a Lightweight Automated AI Pipeline Solve Research-Level Mathematical Problems?

Lve Meng, Weilong Zhao, Yanzhi Zhang +2

Large language models (LLMs) have recently achieved remarkable success in generating rigorous mathematical proofs, with "AI for Math" emerging as a vibrant field of research (Ju et…

cs.AI2025

Population-Evolve: a Parallel Sampling and Evolutionary Method for LLM Math Reasoning

Yanzhi Zhang, Yitong Duan, Zhaoxi Zhang +2

Test-time scaling has emerged as a promising direction for enhancing the reasoning capabilities of Large Language Models in last few years. In this work, we propose Population-Evol…

cs.LG2025

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning

Yanzhi Zhang, Zhaoxi Zhang, Haoxiang Guan +6

Reinforcement learning has emerged as a powerful paradigm for post-training large language models (LLMs) to improve reasoning. Approaches like Reinforcement Learning from Human Fee…

cs.LG2025

EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework

Chen Wang, Lai Wei, Yanzhi Zhang +5

Recent advances in reinforcement learning (RL) have significantly enhanced the reasoning capabilities of large language models (LLMs). Group Relative Policy Optimization (GRPO), a…