works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CV2026

Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently

Simon Roy, Mark Bong, Giovanni Beltrame

The paper studies how to efficiently combine before-and-after satellite images to match natural‑language change queries, comparing attention, state‑space (Mamba), and compressed fu…

cs.RO2026

VISTA: Scale-Aware Visual Navigation via Action History Conditioning

Maeva Guerrier, Koki Kobayashi, Simon Roy +2

Vision Navigation Foundation Models (VNMs) promise end-to-end learned navigation policies capable of zero-shot deployment across diverse embodiments and environments. To maintain g…

cs.AI2026

Latent Goal Prediction from Language for Model-Based Planning

Samuel Barbeau, Simon Roy, Giovanni Beltrame +2

Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals. Visual targets provide precise local gradients but poo…

cs.LG2025

Revisiting the Learning Objectives of Vision-Language Reward Models

Simon Roy, Samuel Barbeau, Giovanni Beltrame +2

Learning generalizable reward functions is a core challenge in embodied intelligence. Recent work leverages contrastive vision language models (VLMs) to obtain dense, domain-agnost…

cs.CL2025

BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices

Euhid Aman, Esteban Carlin, Hsing-Kuo Pao +3

Cross-attention transformers and other multimodal vision-language models excel at grounding and generation; however, their extensive, full-precision backbones make it challenging t…