1 citations · 1 across the 5 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
More Than One Teacher: Adaptive Multi-Guidance Policy Optimization for Diverse Exploration
Xiaoyang Yuan, Yujuan Ding, Yi Bin +5
Reinforcement Learning with Verifiable Rewards (RLVR) is a promising paradigm for enhancing the reasoning ability in Large Language Models (LLMs). However, prevailing methods prima…
cs.CL2025
Explore Briefly, Then Decide: Mitigating LLM Overthinking via Cumulative Entropy Regulation
Yi Bin, Tianyi Jiang, Yujuan Ding +5
Large Language Models (LLMs) have demonstrated remarkable reasoning abilities on complex problems using long Chain-of-Thought (CoT) reasoning. However, they often suffer from overt…
cs.CL2023★ 1 cited
Solving Math Word Problems with Reexamination
Yi Bin, Wenhao Shi, Yujuan Ding +2
Math word problem (MWP) solving aims to understand the descriptive math problem and calculate the result, for which previous efforts are mostly devoted to upgrade different technic…