3 papers
cs.CV2026
Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning
Feiding, Yongkang Zhang, Yuhao Liao +10
Vision--language models (VLMs) are increasingly aligned via Group Relative Policy Optimization (GRPO)-style training. However, relying solely on terminal outcome rewards yields spa…
cs.LG2025
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
Fei Ding, Baiqiao Wang, Zijian Zeng +1
The Group Relative Policy Optimization (GRPO) algorithm has demonstrated considerable success in enhancing the reasoning capabilities of large language models (LLMs), as evidenced…
cs.CL2024
Deep Sparse Latent Feature Models for Knowledge Graph Completion
Haotian Li, Rui Zhang, Lingzhi Wang +6
Recent advances in knowledge graph completion (KGC) have emphasized text-based approaches to navigate the inherent complexities of large-scale knowledge graphs (KGs). While these m…