activity
20242026
most citedERNIE 5.0 Technical Report

2 citations · 2 across the 17 of their papers we have counts for

collaborators

18 papers

cs.AI2026

CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning

Runze Liu, Naibin Gu, Mingxu Ai +4

Continual multimodal instruction tuning requires multimodal large language models to acquire new task abilities sequentially while preserving previously learned knowledge. LoRA-MoE…

cs.CV2026

Harnessing Streaming Video in the Wild

Dingyu Yao, Shuhuan Gu, Qingyi Si +8

Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentary, and embodied robots. An i…

cs.LG2026

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR

Chuanyu Qin, Chenxu Yang, Qingyi Si +3

Reinforcement learning with verifiable rewards (RLVR) improves the ability of large language model, yet headline accuracy gains often conceal a hidden cost: previously solved probl…

cs.LG2026

Co-Evolving Policy Distillation

Naibin Gu, Chenxu Yang, Qingyi Si +7

RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities into a single mode…

cs.LG2026

Near-Future Policy Optimization

Chuanyu Qin, Chenxu Yang, Qingyi Si +6

Reinforcement learning with verifiable rewards (RLVR) has become a core post-training recipe. Introducing suitable off-policy trajectories into on-policy exploration accelerates RL…

cs.CV2026

EasyVideoR1: Easier RL for Video Understanding

Chuanyu Qin, Chenxu Yang, Qingyi Si +6

Reinforcement learning from verifiable rewards (RLVR) has demonstrated remarkable effectiveness in improving the reasoning capabilities of large language models. As models evolve i…