2 papers
cs.LG2026
SPAR: Support-Preserving Action Rectification
Jiaxin Zhao, Weihang Pan, Xun Liang +1
Offline policy improvement faces an inherent conflict between maximizing value and fitting the data distribution. While in-sample weighted regression is stable, it suffers from ove…
cs.CV2025
Group Relative Policy Optimization for Image Captioning
Xu Liang
Image captioning tasks usually use two-stage training to complete model optimization. The first stage uses cross-entropy as the loss function for optimization, and the second stage…