activity
20182026
most citedUpside-Down Reinforcement Learning Can Diverge in Stochastic Environments With Episodic Resets

1 citations · 1 across the 3 of their papers we have counts for

collaborators

8 papers

cs.AI2026

COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami

Tom Zahavy, Shaobo Hou, Thomas Tumiel +16

While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physical art that satisfies both strict geometric constraints and subj…

cs.LG2025

Curious Causality-Seeking Agents Learn Meta Causal World

Zhiyu Zhao, Haoxuan Li, Haifeng Zhang +4

When building a world model, a common assumption is that the environment has a single, unchanging underlying causal rule, like applying Newton's laws to every situation. In reality…

stat.ML2025

On the Convergence and Stability of Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers

Miroslav Štrupl, Oleg Szehr, Francesco Faccio +3

This article provides a rigorous analysis of convergence and stability of Episodic Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning and Online Decision Tran…

cs.LG2025

Upside Down Reinforcement Learning with Policy Generators

Jacopo Di Ventura, Dylan R. Ashley, Vincent Herrmann +2

Upside Down Reinforcement Learning (UDRL) is a promising framework for solving reinforcement learning problems which focuses on learning command-conditioned policies. In this work,…

cs.AI2024

How to Correctly do Semantic Backpropagation on Language-based Agentic Systems

Wenyi Wang, Hisham A. Alyahya, Dylan R. Ashley +4

Language-based agentic systems have shown great promise in recent years, transitioning from solving small-scale research problems to being deployed in challenging real-world tasks.…

stat.ML20221 cited

Upside-Down Reinforcement Learning Can Diverge in Stochastic Environments With Episodic Resets

Miroslav Štrupl, Francesco Faccio, Dylan R. Ashley +2

Upside-Down Reinforcement Learning (UDRL) is an approach for solving RL problems that does not require value functions and uses only supervised learning, where the targets for give…