8 papers
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance
Siyao Song, Cong Ma, Zhihao Cheng +5
Large language models (LLMs) have recently advanced in reasoning when optimized with reinforcement learning (RL) under verifiable rewards. Existing methods primarily rely on outcom…
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
Shiye Lei, Zhihao Cheng, Dacheng Tao
Reinforcement learning (RL) post-training has increasingly demonstrated strong ability to elicit reasoning behaviors in large language models (LLMs). For training efficiency, rollo…
Offline Behavioral Data Selection
Shiye Lei, Zhihao Cheng, Dacheng Tao
Behavioral cloning is a widely adopted approach for offline policy learning from expert demonstrations. However, the large scale of offline behavioral datasets often results in com…
State Diversity Matters in Offline Behavior Distillation
Shiye Lei, Zhihao Cheng, Dacheng Tao
Offline Behavior Distillation (OBD), which condenses massive offline RL data into a compact synthetic behavioral dataset, offers a promising approach for efficient policy training…
EarthSynth: Generating Informative Earth Observation with Diffusion Models
Jiancheng Pan, Shiye Lei, Yuqian Fu +7
Remote sensing image (RSI) interpretation typically faces challenges due to the scarcity of labeled data, which limits the performance of RSI interpretation tasks. To tackle this c…
Revisiting LLM Reasoning via Information Bottleneck
Shiye Lei, Zhihao Cheng, Kai Jia +1
Large language models (LLMs) have recently demonstrated remarkable progress in reasoning capabilities through reinforcement learning with verifiable rewards (RLVR). By leveraging s…