From the 1 of 5 linked papers with an AI index.
5 papers
Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently
Simon Roy, Mark Bong, Giovanni Beltrame
The paper studies how to efficiently combine before-and-after satellite images to match natural‑language change queries, comparing attention, state‑space (Mamba), and compressed fu…
VISTA: Scale-Aware Visual Navigation via Action History Conditioning
Maeva Guerrier, Koki Kobayashi, Simon Roy +2
Vision Navigation Foundation Models (VNMs) promise end-to-end learned navigation policies capable of zero-shot deployment across diverse embodiments and environments. To maintain g…
Latent Goal Prediction from Language for Model-Based Planning
Samuel Barbeau, Simon Roy, Giovanni Beltrame +2
Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals. Visual targets provide precise local gradients but poo…
To Select or not to Select, that is the Question: Distilling Robot Skill Prediction into a Small Ensemble
Haechan Mark Bong, Simon Roy, Euhid Aman +1
As robot fleets become more heterogeneous, including humanoids, rovers, quadrupeds, and drones, selecting the right robot for a task becomes a core systems problem. We study robot…
Revisiting the Learning Objectives of Vision-Language Reward Models
Simon Roy, Samuel Barbeau, Giovanni Beltrame +2
Learning generalizable reward functions is a core challenge in embodied intelligence. Recent work leverages contrastive vision language models (VLMs) to obtain dense, domain-agnost…