3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CL2023★ 3 cited
Iterated Decomposition: Improving Science Q&A by Supervising Reasoning Processes
Justin Reppert, Ben Rachbach, Charlie George +4
Language models (LMs) can perform complex reasoning either end-to-end, with hidden latent state, or compositionally, with transparent intermediate state. Composition offers benefit…
cs.LG2022
Normality-Guided Distributional Reinforcement Learning for Continuous Control
Ju-Seung Byun, Andrew Perrault
Learning a predictive model of the mean return, or value function, plays a critical role in many reinforcement learning algorithms. Distributional reinforcement learning (DRL) has…