4 papers
Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data
Yi Zhao, Aidan Scannell, Wenshuai Zhao +7
Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL). This paper expands the pool of usable data for offline-to-online…
Sequential Causal Discovery with Noisy Language Model Priors
Prakhar Verma, David Arbour, Sunav Choudhary +3
Causal discovery from observational data typically assumes access to complete data and availability of perfect domain experts. In practice, data often arrive in batches, are subjec…
Discrete Codebook World Models for Continuous Control
Aidan Scannell, Mohammadreza Nakhaei, Kalle Kujanpää +4
In reinforcement learning (RL), world models serve as internal simulators, enabling agents to predict environment dynamics and future outcomes in order to make informed decisions.…
Plan*RAG: Efficient Test-Time Planning for Retrieval Augmented Generation
Prakhar Verma, Sukruta Prakash Midigeshi, Gaurav Sinha +3
We introduce Plan*RAG, a novel framework that enables structured multi-hop reasoning in retrieval-augmented generation (RAG) through test-time reasoning plan generation. While exis…