4 citations · 5 across the 27 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification
Yimeng Ye, Shuang Chen, Wenxuan Huang +8
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. Th…
cs.LG2026
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
Zhiqi Yu, Zhangquan Chen, Mengting Liu +2
Reinforcement Learning with Verifiable Rewards (RLVR), particularly GRPO, has become the standard for eliciting LLM reasoning. However, its efficiency in exploration and difficulty…
cs.LG2024
CodingTeachLLM: Empowering LLM's Coding Ability via AST Prior Knowledge
Zhangquan Chen, Chunjiang Liu, Haobin Duan
In this paper, we introduce CodingTeachLLM, a large language model (LLM) designed for coding teaching. Specially, we aim to enhance the coding ability of LLM and lead it to better…