55 citations · 121 across the 11 of their papers we have counts for
4 papers · 1 filter
RLTF: Reinforcement Learning from Unit Test Feedback
Jiate Liu, Yiqin Zhu, Kaiwen Xiao +4
The goal of program synthesis, or code generation, is to generate executable code based on given descriptions. Recently, there has been an increasing number of studies employing re…
Policy Space Diversity for Non-Transitive Games
Jian Yao, Weiming Liu, Haobo Fu +4
Policy-Space Response Oracles (PSRO) is an influential algorithm framework for approximating a Nash Equilibrium (NE) in multi-agent non-transitive games. Many previous studies have…
Future-conditioned Unsupervised Pretraining for Decision Transformer
Zhihui Xie, Zichuan Lin, Deheng Ye +3
Recent research in offline reinforcement learning (RL) has demonstrated that return-conditioned supervised learning is a powerful paradigm for decision-making problems. While promi…
Dynamic Transformers Provide a False Sense of Efficiency
Yiming Chen, Simin Chen, Zexin Li +4
Despite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference. Multi-exit is a mainstream appr…