Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Targeted Exploration via Unified Entropy Control for Reinforcement Learning
Chen Wang, Lai Wei, Yanzhi Zhang +5
Recent advances in reinforcement learning (RL) have improved the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs). However, the widely used…
cs.AI2026
Can a Lightweight Automated AI Pipeline Solve Research-Level Mathematical Problems?
Lve Meng, Weilong Zhao, Yanzhi Zhang +2
Large language models (LLMs) have recently achieved remarkable success in generating rigorous mathematical proofs, with "AI for Math" emerging as a vibrant field of research (Ju et…
cs.AI2025
Population-Evolve: a Parallel Sampling and Evolutionary Method for LLM Math Reasoning
Yanzhi Zhang, Yitong Duan, Zhaoxi Zhang +2
Test-time scaling has emerged as a promising direction for enhancing the reasoning capabilities of Large Language Models in last few years. In this work, we propose Population-Evol…