2 papers
cs.CL2026
Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning
Lei Huang, Xiang Cheng, Chenxiao Zhao +6
Large language models (LLMs) typically receive diverse natural language (NL) feedback through interaction with the environment. However, current reinforcement learning (RL) algorit…
cs.CL2025
We Need Knowledge Distillation for Solving Math Word Problems
Zhenquan Shen, Xinguo Yu, Xiaotian Cheng +2
The enhancement of mathematical capabilities in large language models (LLMs) fosters new developments in mathematics education within primary and secondary schools, particularly as…