2 papers
cs.CV2026
CFPO: Counterfactual Policy Optimization for Multimodal Reasoning
Zhangyuan Yu, Wanran Sun, Guangjing Yang +2
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in multimodal reasoning. However, prevailing reinforcement learning (RL) paradigms lack explicit coun…
cs.CV2026
Improving Medical Visual Reinforcement Fine-Tuning via Perception and Reasoning Augmentation
Guangjing Yang, ZhangYuan Yu, Ziyuan Qin +7
While recent advances in Reinforcement Fine-Tuning (RFT) have shown that rule-based reward schemes can enable effective post-training for large language models, their extension to…