1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Zhiqiang Wang, Pengbin Feng, Yanbin Lin +4
We propose Fuzzy Group Relative Policy Reward (FGRPR), a novel framework that integrates Group Relative Policy Optimization (GRPO) with a fuzzy reward function to enhance learning…