2 papers
cs.CL2026
Language Model Distillation: A Temporal Difference Imitation Learning Perspective
Zishun Yu, Shangzhe Li, Xinhua Zhang
Large language models have led to significant progress across many NLP tasks, although their massive sizes often incur substantial computational costs. Distillation has become a co…
cs.LG2024
Augmenting Offline Reinforcement Learning with State-only Interactions
Shangzhe Li, Xinhua Zhang
Batch offline data have been shown considerably beneficial for reinforcement learning. Their benefit is further amplified by upsampling with generative models. In this paper, we co…