9 papers
Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data
Yi Zhao, Aidan Scannell, Wenshuai Zhao +7
Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL). This paper expands the pool of usable data for offline-to-online…
Sparsely Supervised Diffusion
Wenshuai Zhao, Zhiyuan Li, Yi Zhao +5
Diffusion models have shown remarkable success across a wide range of generative tasks. However, they often suffer from spatially inconsistent generation, arguably due to the inher…
DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing
Qian Cao, Yahui Liu, Wei Bi +6
Reinforcement learning (RL)-based enhancement of large language models (LLMs) often leads to reduced output diversity, undermining their utility in open-ended tasks like creative w…
Slot Attention with Re-Initialization and Self-Distillation
Rongzhen Zhao, Yi Zhao, Juho Kannala +1
Unlike popular solutions based on dense feature maps, Object-Centric Learning (OCL) represents visual scenes as sub-symbolic object-level feature vectors, termed slots, which are h…
Dexterous Robotic Piano Playing at Scale
Le Chen, Yi Zhao, Jan Schneider +7
Endowing robot hands with human-level dexterity has been a long-standing goal in robotics. Bimanual robotic piano playing represents a particularly challenging task: it is high-dim…
Optimistic Multi-Agent Policy Gradient
Wenshuai Zhao, Yi Zhao, Zhiyuan Li +2
*Relative overgeneralization* (RO) occurs in cooperative multi-agent learning tasks when agents converge towards a suboptimal joint policy due to overfitting to suboptimal behavior…