3 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.AI2024★ 3 cited
Zero-shot Imitation Policy via Search in Demonstration Dataset
Federco Malato, Florian Leopold, Andrew Melnik +1
Behavioral cloning uses a dataset of demonstrations to learn a policy. To overcome computationally expensive training procedures and address the policy adaptation problem, we propo…
cs.AI2023★ 1 cited
Towards Solving Fuzzy Tasks with Human Feedback: A Retrospective of the MineRL BASALT 2022 Competition
Stephanie Milani, Anssi Kanervisto, Karolis Ramanauskas +27
To facilitate research in the direction of fine-tuning foundation models from human feedback, we held the MineRL BASALT Competition on Fine-Tuning from Human Feedback at NeurIPS 20…