1 citations · 1 across the 2 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
Kirill Pavlenko, Alexander Golubev, Simon Karasik +1
Group Relative Policy Optimization (GRPO) assigns a single scalar advantage to all tokens in a completion. For structured generations with explicit segments and objectives, this co…
cs.LG2025★ 1 cited
Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning
Alexander Golubev, Maria Trofimova, Sergei Polezhaev +9
Research on applications of reinforcement learning (RL) to large language models has mostly been focused on single-turn problems, such as mathematical reasoning or single-shot code…