Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Gradient Starvation in Binary-Reward GRPO: Why Group-Mean Centering Fails and Why the Simplest Fix Works
Wenhua Nie, Jianan Wu, Junlin Liu +6
Group Relative Policy Optimization (GRPO) is a standard algorithm for reinforcement learning from verifiable rewards, but its group-mean-centered advantage can fail under binary re…
cs.LG2026
The Coupling Tax: How Shared Token Budgets Undermine Visible Chain-of-Thought Under Fixed Output Limits
Wenhua Nie, Junlin Liu, Jianan Wu +5
Chain-of-thought reasoning is often treated as a monotone way to improve language-model accuracy by letting a model think longer. We identify a countervailing effect, the coupling…
cs.LG2023
Enhancing Signed Graph Neural Networks through Curriculum-Based Training
Zeyu Zhang, Lu Li, Xingyu Ji +5
Signed graphs are powerful models for representing complex relations with both positive and negative connections. Recently, Signed Graph Neural Networks (SGNNs) have emerged as pot…