Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Xingjian Zhang, Siwei Wen, Wenjun Wu +1
Large Language Models (LLMs) have made remarkable progress in enhancing step-by-step reasoning through reinforcement learning. However, the Group Relative Policy Optimization (GRPO…
cs.AI2024
Bench-CoE: a Framework for Collaboration of Experts from Benchmark
Yuanshuai Wang, Xingjian Zhang, Jinkun Zhao +5
Large Language Models (LLMs) are key technologies driving intelligent systems to handle multiple tasks. To meet the demands of various tasks, an increasing number of LLMs-driven ex…