behavioral steering 1composable interventions 1decision making 1policy gradient 1reinforcement learning 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
Policy Gradient Steering: Interventions from Behavioral Objectives
Yoann Poupart, Aurélie Beynier, Nicolas Maudet
The paper introduces Policy Gradient Steering (PGS), a reinforcement‑learning based method that uses temporary behavioral objectives to compute removable task vectors for steering…
cs.AI2026
TDHook: A Lightweight Framework for Interpretability
Yoann Poupart
Interpretability of Deep Neural Networks (DNNs) is a growing field driven by the study of vision and language models. Yet, some use cases, like image captioning, or domains like De…
cs.AI2025
Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
Yoann Poupart, Aurélie Beynier, Nicolas Maudet
Multi-Agent Deep Reinforcement Learning (MADRL) was proven efficient in solving complex problems in robotics or games, yet most of the trained models are hard to interpret. While l…