works on

From the 1 of 13 linked papers with an AI index.

activity
20242026
collaborators

13 papers

cs.CV2026

Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction

Lei Yang, Xinze Liu, Dayan Wu +7

The paper introduces a method to detect and correct hallucinated objects in large vision‑language models by identifying a hidden grounding pattern and using a lightweight verifier…

cs.LG2026

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR

Chuanyu Qin, Chenxu Yang, Qingyi Si +3

Reinforcement learning with verifiable rewards (RLVR) improves the ability of large language model, yet headline accuracy gains often conceal a hidden cost: previously solved probl…

cs.CV2026

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

Yuchen Feng, Zhenyu Zhang, Naibin Gu +8

Multimodal large language models (MLLMs) have achieved remarkable progress on various vision-language tasks, yet their visual perception remains limited. Humans, in comparison, per…

cs.CL2026

Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts

Naibin Gu, Zhenyu Zhang, Yuchen Feng +8

Mixture-of-Experts (MoE) models typically fix the number of activated experts at both training and inference. However, real-world deployments often face heterogeneous hardware,…

cs.LG2026

Co-Evolving Policy Distillation

Naibin Gu, Chenxu Yang, Qingyi Si +7

RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities into a single mode…

cs.LG2026

Near-Future Policy Optimization

Chuanyu Qin, Chenxu Yang, Qingyi Si +6

Reinforcement learning with verifiable rewards (RLVR) has become a core post-training recipe. Introducing suitable off-policy trajectories into on-policy exploration accelerates RL…