2 citations · 4 across the 4 of their papers we have counts for
4 papers
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
Zhiheng Xi, Xin Guo, Yang Nan +18
Reinforcement learning (RL) has recently become the core paradigm for aligning and strengthening large language models (LLMs). Yet, applying RL in off-policy settings--where stale…
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
Shaopeng Zhai, Qi Zhang, Tianyi Zhang +7
Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration. We introduce VLA…
Universal Multi-modal Entity Alignment via Iteratively Fusing Modality Similarity Paths
Bolin Zhu, Xiaoze Liu, Xin Mao +4
The objective of Entity Alignment (EA) is to identify equivalent entity pairs from multiple Knowledge Graphs (KGs) and create a more comprehensive and unified KG. The majority of E…
Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games
Dingyang Chen, Yile Li, Qi Zhang
Recent success in cooperative multi-agent reinforcement learning (MARL) relies on centralized training and policy sharing. Centralized training eliminates the issue of non-stationa…