activity
20202026
most citedDREAM Architecture: a Developmental Approach to Open-Ended Learning in Robotics

6 citations · 13 across the 11 of their papers we have counts for

collaborators
Showing cs.AIShow all

9 papers · 1 filter

cs.AI2026

PRISM: Perception Reasoning Interleaved for Sequential Decision Making

Mohamed Salim Aissi, Clemence Grislain, Clement Romac +4

Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception-reasoning-decision gap i…

cs.AI2026

Improving Zero-Shot Offline RL via Behavioral Task Sampling

Nazim Bendib, Nicolas Perrin-Gilbert, Olivier Sigaud

Offline zero-shot reinforcement learning (RL) aims to learn agents that optimize unseen reward functions without additional environment interaction. The standard approach to this p…

cs.AI2025

MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces

Loris Gaven, Thomas Carta, Clément Romac +4

Open-ended learning agents must efficiently prioritize goals in vast possibility spaces, focusing on those that maximize learning progress (LP). When such autotelic exploration is…

cs.AI2024

From Goal-Conditioned to Language-Conditioned Agents via Vision-Language Models

Theo Cachet, Christopher R. Dance, Olivier Sigaud

Vision-language models (VLMs) have tremendous potential for grounding language, and thus enabling language-conditioned agents (LCAs) to perform diverse tasks specified with text. T…

cs.AI20221 cited

Help Me Explore: Minimal Social Interventions for Graph-Based Autotelic Agents

Ahmed Akakzia, Olivier Serris, Olivier Sigaud +1

In the quest for autonomous agents learning open-ended repertoires of skills, most works take a Piagetian perspective: learning trajectories are the results of interactions between…

cs.AI2021

Selection-Expansion: A Unifying Framework for Motion-Planning and Diversity Search Algorithms

Alexandre Chenu, Nicolas Perrin-Gilbert, Stéphane Doncieux +1

Reinforcement learning agents need a reward signal to learn successful policies. When this signal is sparse or the corresponding gradient is deceptive, such agents need a dedicated…