2 papers
cs.CL2024
Into the Unknown: Generating Geospatial Descriptions for New Environments
Tzuf Paz-Argaman, John Palowitch, Sayali Kulkarni +2
Similar to vision-and-language navigation (VLN) tasks that focus on bridging the gap between vision and language for embodied navigation, the new Rendezvous (RVS) task requires rea…
cs.CL2023
Apollo: Zero-shot MultiModal Reasoning with Multiple Experts
Daniela Ben-David, Tzuf Paz-Argaman, Reut Tsarfaty
We propose a modular framework that leverages the expertise of different foundation models over different modalities and domains in order to perform a single, complex, multi-modal…