7 papers
PRISM: Perception Reasoning Interleaved for Sequential Decision Making
Mohamed Salim Aissi, Clemence Grislain, Clement Romac +4
Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception-reasoning-decision gap i…
Online Self-Training for Co-Adaptation in Hierarchical Diffusion Policies
Clemence Grislain, Mathilde Kappel, Olivier Sigaud +1
Hierarchical policies decompose language-conditioned long-horizon robotic manipulation into a high-level planner and a low-level controller. However, effective coordination between…
IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models
Hamed Rahimi, Clemence Grislain, Adrien Jacquet Cretides +2
Improving the effectiveness of human-robot interaction requires social robots to accurately infer human goals through robust intention understanding. This challenge is particularly…
I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
Clemence Grislain, Hamed Rahimi, Olivier Sigaud +1
Language-conditioned robotic manipulation in open-world settings requires not only accurate task execution but also the ability to detect failures for robust deployment in real-wor…
Controlling Intent Expressiveness in Robot Motion with Diffusion Models
Wenli Shi, Clemence Grislain, Olivier Sigaud +1
Legibility of robot motion is critical in human-robot interaction, as it allows humans to quickly infer a robot's intended goal. Although traditional trajectory generation methods…
VIPER: Visual Perception and Explainable Reasoning for Sequential Decision-Making
Mohamed Salim Aissi, Clemence Grislain, Mohamed Chetouani +3
While Large Language Models (LLMs) excel at reasoning on text and Vision-Language Models (VLMs) are highly effective for visual perception, applying those models for visual instruc…