18 citations · 59 across the 28 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Reasoning emerges from constrained inference manifolds in large language models
Yanbiao Ma, Fei Luo, Linfeng Zhang +10
Reasoning in large language models is predominantly evaluated through labeled benchmarks, conflating task performance with the quality of internal inference. Here we study reasonin…
cs.LG2025★ 1 cited
Agentic Entropy-Balanced Policy Optimization
Guanting Dong, Licheng Bao, Zhongyuan Wang +11
Recently, Agentic Reinforcement Learning (Agentic RL) has made significant progress in incentivizing the multi-turn, long-horizon tool-use capabilities of web agents. While mainstr…
cs.LG2025★ 4 cited
Agentic Reinforced Policy Optimization
Guanting Dong, Hangyu Mao, Kai Ma +11
Large-scale reinforcement learning with verifiable rewards (RLVR) has demonstrated its effectiveness in harnessing the potential of large language models (LLMs) for single-turn rea…