Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment
Chenyu Zhou, Qiliang Jiang, Shuning Wu +1
Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattere…
cs.LG2026
Mechanism-Guided Selective Unlearning for RLVR-Induced Reasoning
Chenyu Zhou, Qiliang Jiang, Shuning Wu +1
We propose MAST (Mechanism-Aligned Selective Targeting), a mechanism-guided method for unlearning RLVR-induced reasoning with substantially lower collateral damage than standard fu…