From the 1 of 13 linked papers with an AI index.
13 papers
Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction
Lei Yang, Xinze Liu, Dayan Wu +7
The paper introduces a method to detect and correct hallucinated objects in large vision‑language models by identifying a hidden grounding pattern and using a lightweight verifier…
Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR
Chuanyu Qin, Chenxu Yang, Qingyi Si +3
Reinforcement learning with verifiable rewards (RLVR) improves the ability of large language model, yet headline accuracy gains often conceal a hidden cost: previously solved probl…
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding
Yuchen Feng, Zhenyu Zhang, Naibin Gu +8
Multimodal large language models (MLLMs) have achieved remarkable progress on various vision-language tasks, yet their visual perception remains limited. Humans, in comparison, per…
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
Naibin Gu, Zhenyu Zhang, Yuchen Feng +8
Mixture-of-Experts (MoE) models typically fix the number of activated experts at both training and inference. However, real-world deployments often face heterogeneous hardware,…
Co-Evolving Policy Distillation
Naibin Gu, Chenxu Yang, Qingyi Si +7
RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities into a single mode…
Near-Future Policy Optimization
Chuanyu Qin, Chenxu Yang, Qingyi Si +6
Reinforcement learning with verifiable rewards (RLVR) has become a core post-training recipe. Introducing suitable off-policy trajectories into on-policy exploration accelerates RL…