behavioral steering 1composable interventions 1decision making 1policy gradient 1reinforcement learning 1
From the 1 of 4 linked papers with an AI index.
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
TDHook: A Lightweight Framework for Interpretability
Yoann Poupart
Interpretability of Deep Neural Networks (DNNs) is a growing field driven by the study of vision and language models. Yet, some use cases, like image captioning, or domains like De…
cs.AI2025
Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
Yoann Poupart, Aurélie Beynier, Nicolas Maudet
Multi-Agent Deep Reinforcement Learning (MADRL) was proven efficient in solving complex problems in robotics or games, yet most of the trained models are hard to interpret. While l…
cs.AI2024
Contrastive Sparse Autoencoders for Interpreting Planning of Chess-Playing Agents
Yoann Poupart
AI led chess systems to a superhuman level, yet these systems heavily rely on black-box algorithms. This is unsustainable in ensuring transparency to the end-user, particularly whe…