Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
Hongcheng Wang, Yinuo Huang, Sukai Wang +2
Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers f…
cs.CL2024
PGSO: Prompt-based Generative Sequence Optimization Network for Aspect-based Sentiment Analysis
Hao Dong, Wei Wei
Recently, generative pre-training based models have demonstrated remarkable results on Aspect-based Sentiment Analysis (ABSA) task. However, previous works overemphasize crafting v…