7 papers
Drawback of Enforcing Equivariance and its Compensation via the Lens of Expressive Power
Yuzhu Chen, Tian Qin, Xinmei Tian +2
Equivariant neural networks encode the intrinsic symmetry of data as an inductive bias, which has achieved impressive performance in wide domains. However, the understanding to the…
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
Shiye Lei, Zhihao Cheng, Dacheng Tao
Reinforcement learning (RL) post-training has increasingly demonstrated strong ability to elicit reasoning behaviors in large language models (LLMs). For training efficiency, rollo…
Offline Behavioral Data Selection
Shiye Lei, Zhihao Cheng, Dacheng Tao
Behavioral cloning is a widely adopted approach for offline policy learning from expert demonstrations. However, the large scale of offline behavioral datasets often results in com…
State Diversity Matters in Offline Behavior Distillation
Shiye Lei, Zhihao Cheng, Dacheng Tao
Offline Behavior Distillation (OBD), which condenses massive offline RL data into a compact synthetic behavioral dataset, offers a promising approach for efficient policy training…
Revisiting LLM Reasoning via Information Bottleneck
Shiye Lei, Zhihao Cheng, Kai Jia +1
Large language models (LLMs) have recently demonstrated remarkable progress in reasoning capabilities through reinforcement learning with verifiable rewards (RLVR). By leveraging s…
Image Captions are Natural Prompts for Text-to-Image Models
Shiye Lei, Hao Chen, Sen Zhang +2
With the rapid development of Artificial Intelligence Generated Content (AIGC), it has become a common practice to train models on synthetic data due to data-scarcity and privacy l…