5 papers
Latent Goal Prediction from Language for Model-Based Planning
Samuel Barbeau, Simon Roy, Giovanni Beltrame +2
Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals. Visual targets provide precise local gradients but poo…
TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses
Sahar Dastani, Ali Bahri, Gustavo Adolfo Vargas Hakim +7
State Space Models (SSMs) have emerged as efficient alternatives to Vision Transformers (ViTs), with VMamba standing out as a pioneering architecture designed for vision tasks. How…
Revisiting the Learning Objectives of Vision-Language Reward Models
Simon Roy, Samuel Barbeau, Giovanni Beltrame +2
Learning generalizable reward functions is a core challenge in embodied intelligence. Recent work leverages contrastive vision language models (VLMs) to obtain dense, domain-agnost…
DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
Pedram Fekri, Majid Roshanfar, Samuel Barbeau +5
Cardiac catheterization remains a cornerstone of minimally invasive interventions, yet it continues to rely heavily on manual operation. Despite advances in robotic platforms, exis…
CTA: Cross-Task Alignment for Better Test Time Training
Samuel Barbeau, Pedram Fekri, David Osowiechi +4
Deep learning models have demonstrated exceptional performance across a wide range of computer vision tasks. However, their performance often degrades significantly when faced with…