8 papers
PRISM: Perception Reasoning Interleaved for Sequential Decision Making
Mohamed Salim Aissi, Clemence Grislain, Clement Romac +4
Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception-reasoning-decision gap i…
MODIP: Efficient Model-Based Optimization for Diffusion Policies
Zakariae El Asri, Philippe Gratias-Quiquandon, Nicolas Thome +1
Diffusion policies (DPs) have emerged as expressive policy representations for robot learning, often used with imitation learning methods such as behavioral cloning (BC). However,…
Latent Goal Prediction from Language for Model-Based Planning
Samuel Barbeau, Simon Roy, Giovanni Beltrame +2
Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals. Visual targets provide precise local gradients but poo…
Boosting Visual Instruction Tuning with Self-Supervised Guidance
Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc +2
Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-grained visual reasoning. Rece…
Revisiting the Learning Objectives of Vision-Language Reward Models
Simon Roy, Samuel Barbeau, Giovanni Beltrame +2
Learning generalizable reward functions is a core challenge in embodied intelligence. Recent work leverages contrastive vision language models (VLMs) to obtain dense, domain-agnost…
VIPER: Visual Perception and Explainable Reasoning for Sequential Decision-Making
Mohamed Salim Aissi, Clemence Grislain, Mohamed Chetouani +3
While Large Language Models (LLMs) excel at reasoning on text and Vision-Language Models (VLMs) are highly effective for visual perception, applying those models for visual instruc…