3 papers
cs.CL2025
Language Model Distillation: A Temporal Difference Imitation Learning Perspective
Zishun Yu, Shangzhe Li, Xinhua Zhang
Large language models have led to significant progress across many NLP tasks, although their massive sizes often incur substantial computational costs. Distillation has become a co…
cs.MA2024
Towards Efficient Collaboration via Graph Modeling in Reinforcement Learning
Wenzhe Fan, Zishun Yu, Chengdong Ma +3
In multi-agent reinforcement learning, a commonly considered paradigm is centralized training with decentralized execution. However, in this framework, decentralized execution rest…
cs.CL2023
-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis
Zishun Yu, Yunzhe Tao, Liyu Chen +2
Program synthesis aims to create accurate, executable programs from problem specifications, specifically from natural language descriptions in our context. Recent studies have leve…