3 papers
cs.AI2026
FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
Aniri, Chen Yilin, Jinhe Bi +10
Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures. However,…
cs.CV2026
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Aniri, Jinhe Bi, Peng Liao +5
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw pr…
cs.AI2026
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
Jinhe Bi, Chennan Zhou, Zengjie Jin +10
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories…