most citedEDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

1 citations · 2 across the 6 of their papers we have counts for

collaborators

9 papers

cs.LG2025

RLLaVA: An RL-central Framework for Language and Vision Assistants

Lei Zhao, Zihao Ma, Boyu Lin +3

We present an RL-central framework for Language and Vision Assistants (RLLaVA) with its formulation of Markov decision process (MDP). RLLaVA decouples RL algorithmic logic from mod…

cs.AI2025

Tool-Augmented Hybrid Ensemble Reasoning with Distillation for Bilingual Mathematical Problem Solving

Peiqing Lu, Yuan Zhang, Haoyun Zhang +3

Bilingual mathematical problem solving needs a clear link between language reasoning and symbolic calculation. Large language models often handle language well but are weak in accu…

cs.AI20251 cited

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Xingjian Zhang, Siwei Wen, Wenjun Wu +1

Large Language Models (LLMs) have made remarkable progress in enhancing step-by-step reasoning through reinforcement learning. However, the Group Relative Policy Optimization (GRPO…

cs.CL2025

CoE-Ops: Collaboration of LLM-based Experts for AIOps Question-Answering

Jinkun Zhao, Yuanshuai Wang, Xingjian Zhang +6

With the rapid evolution of artificial intelligence, AIOps has emerged as a prominent paradigm in DevOps. Lots of work has been proposed to improve the performance of different AIO…

cs.CV2025

VoQA: Visual-only Question Answering

Jianing An, Luyang Jiang, Jie Luo +2

Visual understanding requires interpreting both natural scenes and the textual information that appears within them, motivating tasks such as Visual Question Answering (VQA). Howev…

cs.CV2025

TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Xingjian Zhang, Siwei Wen, Wenjun Wu +1

Recently, improving the reasoning ability of large multimodal models (LMMs) through reinforcement learning has made great progress. However, most existing works are based on highly…