3 papers
cs.LG2026
Learning-Zone Energy: Online Data Selection for Efficient RL Post-Training
Peng Cui, Boyao Yang, Jun Zhu
Reinforcement Learning (RL) post-training has emerged as the dominant paradigm for eliciting mathematical reasoning in Large Language Models (LLMs), yet prevailing techniques such…
cs.CV2026
SkyNative: A Native Multimodal Architecture for Remote Sensing Vision-Language Understanding
Xiao Yang, Ronghao Fu, Zhiwen Lin +10
Remote sensing vision-language models (RS-VLMs) commonly employ a pretrained vision encoder and a projection module to map image features into the token space of a large language m…
cs.LG2026
Ranking-Aware Calibration for Reliable Multimodal Reinforcement Learning
Peng Cui, Boyao Yang, Jun Zhu
Reinforcement learning post-training has substantially improved the reasoning accuracy of vision-language models, yet the resulting policies remain poorly calibrated. Terminal corr…