collaborators

6 papers

cs.LG2026

Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning

Thomas Carta, Clément Romac, Thomas Wolf +3

Recent works successfully leveraged Large Language Models' (LLM) abilities to capture abstract knowledge about world's physics to solve decision-making problems. Yet, the alignment…

cs.AI2026

WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making

Guillaume Levy, Cedric Colas, Pierre-Yves Oudeyer +2

Large Language Models (LLMs) possess general world knowledge but often struggle to generate precise predictions in structured, domain-specific contexts such as simulations. These l…

cs.LG2026

SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling

Loris Gaven, Clement Romac, Thomas Carta +3

The past years have seen Large Language Models (LLMs) strive not only as generative models but also as agents solving textual sequential decision-making tasks. When facing complex…

cs.LG2025

Reinforcement Learning for Aligning Large Language Models Agents with Interactive Environments: Quantifying and Mitigating Prompt Overfitting

Mohamed Salim Aissi, Clement Romac, Thomas Carta +5

Reinforcement learning (RL) is a promising approach for aligning large language models (LLMs) knowledge with sequential decision-making tasks. However, few studies have thoroughly…

cs.LG2025

HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents

Thomas Carta, Clément Romac, Loris Gaven +3

We study goal-conditioned reinforcement learning in partially observable environments with sparse rewards and large, structured goal spaces. In such settings, complex goals often r…

cs.AI2025

MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces

Loris Gaven, Thomas Carta, Clément Romac +4

Open-ended learning agents must efficiently prioritize goals in vast possibility spaces, focusing on those that maximize learning progress (LP). When such autotelic exploration is…