2 papers
cs.LG2026
On the Plasticity and Stability for Post-Training Large Language Models
Wenwen Qiang, Ziyin Gu, Jiahuan Zhou +4
Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability…
cs.CV2024
Intriguing Property and Counterfactual Explanation of GAN for Remote Sensing Image Generation
Xingzhe Su, Wenwen Qiang, Jie Hu +3
Generative adversarial networks (GANs) have achieved remarkable progress in the natural image field. However, when applying GANs in the remote sensing (RS) image generation task, a…