most citedKlear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning

1 citations · 1 across the 12 of their papers we have counts for

collaborators

16 papers

cs.CL2026

DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing

Hongzhi Zhang, Yuanze Hu, Tinghai Zhang +9

The evolution of Large Language Models (LLMs) towards autonomous agents has catalyzed progress in Deep Research. While retrieval capabilities are well-benchmarked, the post-retriev…

cs.AI2025

Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling

Qi Wang, Hongzhi Zhang, Jia Fu +12

Despite the proliferation of powerful agentic models, the lack of critical post-training details hinders the development of strong counterparts in the open-source community. In thi…

cs.SE20251 cited

Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning

Jia Fu, Xinyu Yang, Hongzhi Zhang +5

Precise, correct feedback is crucial for effectively training large language models (LLMs) in code reinforcement learning. However, synthesizing high-quality test cases remains a p…

cs.CV2025

AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning

Shihao Yuan, Yahui Liu, Yang Yue +5

Inspired by the success of reinforcement learning (RL) in refining large language models (LLMs), we propose AR-GRPO, an approach to integrate online RL training into autoregressive…

cs.AI2025

Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning

Xingguang Ji, Yahui Liu, Qi Wang +7

We introduce our Leanabell-Prover-V2, a 7B large language models (LLMs) that can produce formal theorem proofs in Lean 4, with verifier-integrated Long Chain-of-Thoughts (CoT). Fol…

cs.CL2025

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Hongzhi Zhang, Jia Fu, Jingyuan Zhang +4

Reinforcement learning (RL) for large language models is an energy-intensive endeavor: training can be unstable, and the policy may gradually drift away from its pretrained weights…