activity
20152025
most citedReinforced Self-Training (ReST) for Language Modeling

18 citations · 57 across the 9 of their papers we have counts for

collaborators

16 papers

cs.RO2025★ 1 cited

Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

Gemini Robotics Team, Abbas Abdolmaleki, Saminda Abeyruwan +169

General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the G…

cs.LG2025

Vision-Language Model Dialog Games for Self-Improvement

Ksenia Konyushkova, Christos Kaplanis, Serkan Cabi +1

The increasing demand for high-quality, diverse training data poses a significant bottleneck in advancing vision-language models (VLMs). This paper presents VLM Dialog Games, a nov…

cs.CL2023★ 18 cited

Reinforced Self-Training (ReST) for Language Modeling

Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan +11

Reinforcement learning from human feedback (RLHF) can improve the quality of large language model's (LLM) outputs by aligning them with human preferences. We propose a simple algor…

cs.LG2023

: Policy Representations with Successor Features

Gianluca Scarpellini, Ksenia Konyushkova, Claudio Fantacci +3

This paper describes , a method for representing behaviors of black box policies as feature vectors. The policy representations capture how the statistics of founda…

cs.CV2023★ 11 cited

Vision-Language Models as Success Detectors

Yuqing Du, Ksenia Konyushkova, Misha Denil +5

Detecting successful behaviour is crucial for training intelligent agents. As such, generalisable reward models are a prerequisite for agents that can learn to generalise their beh…

cs.LG2022★ 6 cited

Retrieval-Augmented Reinforcement Learning

Anirudh Goyal, Abram L. Friesen, Andrea Banino +13

Most deep reinforcement learning (RL) algorithms distill experience into parametric behavior policies or value functions via gradient updates. While effective, this approach has se…